Machine learning in automated text categorization
Explore this paper's citation graph
Summary
This survey discusses the main approaches to text categorization that fall within the machine learning paradigm and discusses in detail issues pertaining to three different problems, namely, document representation, classifier construction, and classifier evaluation.
- Type
- review
- Published
- 2001-10-26
- Cited by
- 9,107
- References
- 198
- Access
- Open access
- OpenAlex
- https://openalex.org/W2118020653
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:3091
Keywords
Computer science, Categorization, Software portability, Artificial intelligence, Classifier (UML)
References
- Categorisation by Context
- Using a Bayesian Network Induction Approach for Text Categorization
- ATTICS: A Software Platform for Online Text Classification (poster abstract).
- Classification Trees for Document Routing, A Report on the TREC Experiment
- Boosting Applied toe Word Sense Disambiguation
- A probabilistic model of dictionary based automatic indexing
- Active Learning with Committees for Text Categorization
- Document Retrieval Systems
- Integrating linguistic resources in a uniform way for Text classification tasks
- Learning Rules that Classify E-Mail
- The EURATOM automatic indexing project
- Readings in information retrieval.
- A Study of Approaches to Hypertext Categorization
- Research and Development in Information Retrieval
- Exploiting Hierarchy in Text Categorization
- Feature Selection in SVM Text Categorization
- Foundations of Statistical Natural Language Processing
- A Decision Theory Approach to Optimal Automatic Indexing
- Classification of text documents
- Employing EM and Pool-Based Active Learning for Text Classification
Cited by
- A modular architecture for systematic text categorisation
- Refining Search Queries From Examples Using Boolean Expressions and Latent Semantic Analysis
- Multi-facet Classification of E-mails in a Helpdesk Scenario
- Finding Communities of Practice from User Profiles Based on Folksonomies
- Espaces vectoriels sémantiques : enrichissement et interprétation de requêtes dans un système d'information distribué et hétérogène. (Semantic Vector Spaces: Query Enrichment and Interpretation in a Distributed and Heterogeneous Information System)
- Threshold Optimization with a Small Number of Samples in Adaptive Information Filtering
- Sequence Learning from Data with Multiple Labels
- Natural Language Processing and Text Mining
- Does a New Simple Gaussian Weighting Approach Perform Well in Text Categorization?
- A Comparative study on text categorization
- A Survey of Retrieval Strategies for OCR Text Collections
- Domain Parking Recognizer: an experimental study on web content categorization
- Semantic and Bayesian Profiling Services for Textual Resource Retrieval
- Automating Creation of Hierarchical Faceted Metadata Structures
- Conglomerate Industry Spanning
- An Evaluation of Bag-of-Concepts Representations in Automatic Text Classification
- Informations morpho-syntaxiques et adaptation thématique pour améliorer la reconnaissance de la parole
- Managing Content with Automatic Document Classification
- Profiling Users to Perform Contextual Advertising
- Language identification of person names using CF-IOF based weighing function
Related papers
- Feature Selection in Text Categorization
- Automatic text categorization for patent data
- Blog Categorization Exploiting Domain Dictionary and Dynamically Estimated Domains of Unknown Words
- Study on Improved CHI for feature selection in Chinese text categorization
- Modern Text Categorization Technology Analyse
- On the strength of hyperclique patterns for text categorization
- k-NN Text Categorization Method Based on Transferable Belief Model
- Realization of Text Categorization for Small-Scaled Dataset
- Documents Categorization in Multilingual Environment