Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift
Explore this paper's citation graph
Summary
A large-scale benchmark of existing state-of-the-art methods on classification problems and the effect of dataset shift on accuracy and calibration is presented, finding that traditional post-hoc calibration does indeed fall short, as do several other previous methods.
- Type
- preprint
- Published
- 2019-06-06
- Cited by
- 2,394
- References
- 58
- Access
- Open access
- OpenAlex
- https://openalex.org/W2948194985
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:174803437
Keywords
Machine learning, Benchmark (surveying), Computer science, Artificial intelligence, Calibration
References
- Training Very Deep Networks
- One billion word benchmark for measuring progress in statistical language modeling
- The Comparison and Evaluation of Forecasters.
- Probabilistic Outputs for Support vector Machines and Comparisons to Regularized Likelihood Methods
- Probabilistic Backpropagation for Scalable Learning of Bayesian Neural Networks
- Strictly Proper Scoring Rules, Prediction, and Estimation
- Long Short-Term Memory
- VERIFICATION OF FORECASTS EXPRESSED IN TERMS OF PROBABILITY
- Semi-supervised Learning with Deep Generative Models
- ImageNet: A large-scale hierarchical image database
- Practical Variational Inference for Neural Networks
- Gradient-based learning applied to document recognition
- Dataset Shift in Machine Learning
- Weight Uncertainty in Neural Network
- Reliability, Sufficiency, and the Decomposition of Proper Scores
- Bayesian Learning via Stochastic Gradient Langevin Dynamics
- Deep Residual Learning for Image Recognition
- Obtaining Well Calibrated Probabilities Using Bayesian Binning
- Reading Digits in Natural Images with Unsupervised Feature Learning
- End to End Learning for Self-Driving Cars
Cited by
- A Roadmap for Robust End-to-End Alignment
- Ensemble Distribution Distillation
- Unified Probabilistic Deep Continual Learning through Generative Replay and Open Set Recognition
- Evaluating Scalable Bayesian Deep Learning Methods for Robust Computer Vision
- Likelihood Ratios for Out-of-Distribution Detection
- Open Set Recognition Through Deep Neural Network Uncertainty: Does Out-of-Distribution Detection Require Generative Classifiers?
- Expert-validated estimation of diagnostic uncertainty for deep neural networks in diabetic retinopathy detection
- Uncertainty Quantification with Statistical Guarantees in End-to-End Autonomous Driving Control
- Training Data Distribution Search with Ensemble Active Learning
- BANANAS: Bayesian Optimization with Neural Architectures for Neural Architecture Search
- Detecting Extrapolation with Local Ensembles
- We Know Where We Don't Know: 3D Bayesian CNNs for Uncertainty Quantification of Binary Segmentations for Material Simulations
- Deep learning analyses of synthetic spectral libraries with an application to the Gaia-ESO database
- Deep Verifier Networks: Verification of Deep Discriminative Models with Deep Generative Models
- A Novel Unsupervised Post-Processing Calibration Method for DNNS with Robustness to Domain Shift
- Parameters Estimation for the Cosmic Microwave Background with Bayesian Neural Networks
- Deep Ensembles: A Loss Landscape Perspective
- Individual predictions matter: Assessing the effect of data ordering in training fine-tuned CNNs for medical imaging
- AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty
- Bayesian Variational Autoencoders for Unsupervised Out-of-Distribution Detection
Related papers
- Theoretical Analysis of the Benchmark for Choosing Manipulative Instruments of Monetary Policies
- Exploring disk performance benchmarks
- Solutions to the Third Benchmark Control Problem
- A Benchmark Characterization of the EEMBC Benchmark Suite
- The Performance Validation of Linear Programming Algorithm Based on Integrated Benchmark
- An empirical assessment of Bellon's clone benchmark
- Non-probabilistic credible set model for structural uncertainty quantification