A Comparative Study on Feature Selection in Text Categorization
Explore this paper's citation graph
Summary
DF thresholding, the simplest method with the lowest cost in computation, can be reliably used instead of IG or CHI when the computation of these measures are too expensive, and strong correlations between the DF, IG and CHI values of a term are found.
- Type
- article
- Published
- 1997-07-08
- Cited by
- 5,820
- References
- 28
- OpenAlex
- https://openalex.org/W2435251607
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:5083193
Keywords
Feature selection, Categorization, Text categorization, Computer science, Artificial intelligence
References
- AIR/X - A rule-based multistage indexing system for Iarge subject fields
- Automatic Text Processing: The Transformation, Analysis, and Retrieval of Information by Computer
- Toward Optimal Feature Selection
- Context-sensitive learning methods for text categorization
- Noise reduction in a statistical approach to text categorization
- An example-based mapping method for text categorization and retrieval
- Automatic indexing based on Bayesian inference networks
- Towards language independent automated learning of text categorization models
- Transmission of Information
- A comparison of classifiers and document representations for the routing problem
- Training algorithms for linear text classifiers
- Trading MIPS and memory for knowledge engineering
- Using Corpus Statistics to Remove Redundant Words in Text Categorization
- Word Association Norms, Mutual Information, and Lexicography
- Accurate Methods for the Statistics of Surprise and Coincidence
- A comparison of two learning algorithms for text categorization
- Indexing by Latent Semantic Analysis
- OHSUMED: an interactive retrieval evaluation and new large test collection for research
- A neural network approach to topic spotting
- Expert network: effective and efficient learning from human decisions in text categorization and retrieval
Cited by
- Combining Text and Linguistic Document Representations for Authorship Attribution
- A modular architecture for systematic text categorisation
- Apprentissage automatique pour l'extraction de caractéristiques : application au partitionnement de documents, au résumé automatique et au filtrage collaboratif
- Eksploracja danych - przegląd dostępnych metod i dziedzin zastosowań
- GU METRIC - A New Feature Selection Algorithm for Text Categorization
- Text mining and IRT for psychiatric and psychological assessment
- Feature Selection Using Linear Support Vector Machines
- Detecting Spammers on Twitter
- Natural Language Processing and Text Mining
- Does a New Simple Gaussian Weighting Approach Perform Well in Text Categorization?
- A Comparative study on text categorization
- Text Passage Classification Using Supervised Learning
- Approaches to Feature Selection for Document Categorization
- Mining Personal Data Collections to Discover Categories and Category Labels
- An Evaluation of Bag-of-Concepts Representations in Automatic Text Classification
- The Robert Gordon University at the Opinion Retrieval Task of the 2007 TREC Blog Track
- Managing Content with Automatic Document Classification
- Unsupervised categorisation approaches for technical support automated agents
- Stuctured Queries for Legal Search.
- Feature-based transfer learning with real-world applications
Related papers
- Inductive learning algorithms and representations for text categorization
- A vector space model for automatic indexing
- The Nature of Statistical Learning Theory
- LIBSVM: A library for support vector machines
- Indexing by Latent Semantic Analysis
- C4.5: Programs for Machine Learning
- Support-Vector Networks
- An introduction to variable and feature selection
- Machine learning in automated text categorization