An Extensive Empirical Study of Feature Selection Metrics for Text Classification
Explore this paper's citation graph
Summary
An empirical comparison of twelve feature selection methods evaluated on a benchmark of 229 text classification problem instances, revealing that a new feature selection metric, called 'Bi-Normal Separation' (BNS), outperformed the others by a substantial margin in most situations and was the top single choice for all goals except precision.
- Type
- article
- Published
- 2003-03-01
- Cited by
- 3,005
- References
- 15
- OpenAlex
- https://openalex.org/W3005526988
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:809191
Keywords
Feature selection, Computer science, Artificial intelligence, Selection (genetic algorithm), Feature (linguistics)
References
- Feature Selection for Unbalanced Class Distribution and Naive Bayes
- A comparison of event models for naive bayes text classification
- Learning with Kernels: support vector machines, regularization, optimization, and beyond
- Text Categorization with Support Vector Machines: Learning with Many Relevant Features
- Benchmarking attribute selection techniques for data mining
- What is the best index of detectability?
- The Robustness of the "Binormal" Assumptions Used in Fitting ROC Curves
- A re-examination of text categorization methods
- Wrappers for Feature Subset Selection
- Tests of a statistical explanation of the rank-frequency relation for words in written English.
- A Comparative Study on Feature Selection in Text Categorization
- Centroid-Based Document Classifica tion : Analysis & Exper imental Results ∗
- Centroid-Based Document Classification: Analysis and Experimental Results
- Ììì Öûûò Ë Blockinöö Óóóòòòö Áòøøöòòøøóòòð
Cited by
- What Matters Most? How Tone in Initial Public Offering Filings and Pre-IPO News Influences Stock Market Returns
- Apprentissage automatique pour l'extraction de caractéristiques : application au partitionnement de documents, au résumé automatique et au filtrage collaboratif
- GU METRIC - A New Feature Selection Algorithm for Text Categorization
- Text mining and IRT for psychiatric and psychological assessment
- Context-based Term Disambiguation in Biomedical Literature
- Estimation of Discriminative Feature Subset Using Community Modularity
- A Novel Feature Selection Technique for Text Classification Using Naïve Bayes
- Statistical Approaches for Gene Selection, Hub Gene Identification and Module Interaction in Gene Co-Expression Network Analysis: An Application to Aluminum Stress in Soybean (Glycine max L.)
- Supervised Extraction of Diagnosis Codes from EMRs: Role of Feature Selection, Data Selection, and Probabilistic Thresholding
- A framework for feature selection in high-dimensional domains
- Advances in Machine Learning Based Text Categorization
- Machine Learning for Natural Language Processing
- Reconnaissance de catégories d'objets et d'instances d'objets à l'aide de représentations locales. (Local Feature Based Object Categories and Object Instances Recognition)
- Graph-based Semi-supervised Learning: Realizing Pointwise Smoothness Probabilistically
- Using Closed Captions and Visual Features to Classify Movies by Genre
- Feature Selection for the Classification of Large Document Collections
- Learning from communication data: language in electronic business negotiations
- KACST Arabic Text Classification Project: Overview and Preliminary Results
- Discriminative learning and spanning tree algorithms for dependency parsing
- A Feature Ranking Algorithm in Pragmatic Quality Factor Model for Software Quality Assessment
Related papers
- A Comparative Study on Feature Selection in Text Categorization
- Machine learning in automated text categorization
- A Feature Selection and Classification Technique for Text Categorization
- Improving Text Categorization by Multicriteria Feature Selection
- A New Text Categorization Technique Using Distributional Clustering and Learning Logic
- The research progress of Text Classification Techniques
- Approaches to Feature Selection for Document Categorization
- Evaluation of text classification techniques for inappropriate web content blocking
- Web document classification techniques