A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection
Explore this paper's citation graph
Summary
The results indicate that for real-word datasets similar to the authors', the best method to use for model selection is ten fold stratified cross validation even if computation power allows using more folds.
- Type
- article
- Published
- 1995-08-20
- Cited by
- 13,926
- References
- 26
- Access
- Open access
- OpenAlex
- https://openalex.org/W1680392829
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:2702042
Keywords
Cross-validation, Naive Bayes classifier, Model selection, Computer science, Classifier (UML)
References
- Stacked generalization
- Decision Tree Pruning: Biased or Optimal?
- MLC++, A Machine Learning Library in C++.
- Classification and regression trees
- An Analysis of Bayesian Classifiers
- Estimating the Accuracy of Learned Concepts
- Bootstrap Techniques for Error Estimation
- Estimating the Error Rate of a Prediction Rule: Improvement on Cross-Validation
- Small Sample Error Rate Estimation for k-NN Classifiers
- Submodel selection and evaluation in regression. The X-random case
- C4.5: Programs for Machine Learning
- On the Distributional Properties of Model Selection Criteria
- Neural networks
- OFF-TRAINING SET ERROR AND A PRIORI DISTINCTIONS BETWEEN LEARNING ALGORITHMS
- The Relationship Between PAC, the Statistical Physics Framework, the Bayesian Framework, and the VC Framework
- A Conservation Law for Generalization Performance
Cited by
- A modular architecture for systematic text categorisation
- Software defect prediction using static code metrics : formulating a methodology
- Social networks and collaborative filtering for large-scale requirements elicitation
- A New Metric-Based Approach to Model Selection
- Extraction de couples nom-verbe sémantiquement liés : une technique symbolique automatique
- Machine Learning for First-Order Theorem Proving
- Tense and Mood Decision with Similarity Search in Japanese to Spanish Machine Translation
- Multiattribute Loss Aversion and Reference Dependence: Evidence from the Performing Arts Industry
- A Study of Early Stopping and Model Selection Applied to the Papermaking Industry
- Building bayesian networks from data: a constraint-based approach
- Tree-Based Credal Networks for Classification
- Technologies for Mobile ITS Applications and Safer Driving
- Tolerance to missing data using a likelihood ratio based classifier for computer-aided classification of breast cancer
- Automatic segmentation of diatom images for classification
- Simple point-of-care risk stratification in acute coronary syndromes: the AMIS model
- Multinomial nonparametric predictive inference : selection, classification and subcategory data
- Predicting variations of perceptual performance across individuals from neural activity using pattern classifiers
- On-Line New Event Detection using Single Pass Clustering
- To aggregate or not to aggregate high-dimensional classifiers
- Modulation of anticipatory postural activity for multiple conditions of a whole-body pointing task.
Related papers
- Classification and Regression Trees
- Bagging Predictors
- Random Forests
- An introduction to ROC analysis
- The Nature of Statistical Learning Theory
- LIBSVM: A library for support vector machines
- Induction of Decision Trees
- Statistical learning theory
- Regression Shrinkage and Selection via the Lasso
- The WEKA data mining software: an update