A comparison of event models for naive bayes text classification
Explore this paper's citation graph
Summary
It is found that the multi-variate Bernoulli performs well with small vocabulary sizes, but that the multinomial performs usually performs even better at larger vocabulary sizes--providing on average a 27% reduction in error over the multi -variateBernoulli model at any vocabulary size.
- Type
- article
- Published
- 1998-01-01
- Cited by
- 4,430
- References
- 37
- OpenAlex
- https://openalex.org/W1550206324
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:7311285
Keywords
Computer science, Artificial intelligence, Naive Bayes classifier, Perplexity, Vocabulary
References
- A New Probabilistic Model of Text Classification and Retrieval
- Active Learning with Committees for Text Categorization
- Learning Limited Dependence Bayesian Classifiers
- Information Theory: 1948-1998 - Guest Editorial
- Some Issues in the Automatic Classification of U.S. Patents Working Notes for the AAAI-98 Workshop on Learning for Text Categorization
- Improving Text Classification by Shrinkage in a Hierarchy of Classes
- Text Categorization with Support Vector Machines: Learning with Many Relevant Features
- Hierarchically Classifying Documents Using Very Few Words
- An Analysis of Bayesian Classifiers
- A Bayesian Approach to Filtering Junk E-Mail
- Molecular manipulations of the multidrug transporter: a new role for transgenic mice 1
- Translocation of DNA across bacterial membranes.
- Document Classification by Machine:Theory and Practice
- Combining classifiers in text categorization
- Towards an understanding of the genetics of bacterial metal resistance.
- Relevance weighting of search terms
- Bioenergetic aspects of the translocation of macromolecules across bacterial membranes.
- The regulation of genetic competence in Bacillus subtilis
- A sequential algorithm for training text classifiers
- A Probabilistic Analysis of the Rocchio Algorithm with TFIDF for Text Categorization
Cited by
- Tracking Conversational Context for Machine Mediation of Human Discourse
- Learning to Identify Unexpected Instances in the Test Set
- Approaches to Feature Selection for Document Categorization
- A ME Model Based on Feature Template for Chinese Text Categorization
- A Multiagent Architecture for Information Retrieval in Distributed and Heterogeneous Data Sources
- Hypergraph-Based Anomaly Detection of High-Dimensional Co-Occurrences
- Consensus policies to solve bioinformatic problems through Bayesian network classifiers and estimation of distribution algorithms
- Profiler-2000: Attacking the Insider Threat
- A Novel Feature Selection Technique for Text Classification Using Naïve Bayes
- SegAuth: A Segment-based Approach to Behavioral Biometric Authentication
- Attribute Construction for E-Mail Foldering by Using Wrappered Forward Greedy Search
- A framework for feature selection in high-dimensional domains
- Machine Learning for Natural Language Processing
- TPN 2 : Using positive-only learning to deal with the heterogeneity of labeled and unlabeled data
- Corpus Based Unsupervised Labeling of Documents
- Text Classification through Time - Efficient Label Propagation in Time-Based Graphs
- Fighting phishing at the user interface
- Robust statistical techniques for the categorization of images using associated text
- Yucca Mountain licensing support network archive assistant.
- Automatic Extraction of ICD-O-3 Primary Sites from Cancer Pathology Reports
Related papers
- Random Forests
- A Comparative Study on Feature Selection in Text Categorization
- Inductive learning algorithms and representations for text categorization
- The Nature of Statistical Learning Theory
- On the Optimality of the Simple Bayesian Classifier under Zero-One Loss
- Tackling the Poor Assumptions of Naive Bayes Text Classifiers
- C4.5: Programs for Machine Learning
- Support-Vector Networks
- Machine learning in automated text categorization