Sharpness-Aware Minimization for Efficiently Improving Generalization
Explore this paper's citation graph
Summary
This work introduces a novel, effective procedure for simultaneously minimizing loss value and loss sharpness, Sharpness-Aware Minimization (SAM), which improves model generalization across a variety of benchmark datasets and models, yielding novel state-of-the-art performance for several.
- Type
- preprint
- Published
- 2020-10-03
- Cited by
- 2,067
- References
- 67
- Access
- Open access
- OpenAlex
- https://openalex.org/W3091401866
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:222134093
Keywords
Robustness (evolution), Generalization, Computer science, Benchmark (surveying), Minification
References
- The conference paper
- PAC-Bayesian model averaging
- Dropout: a simple way to prevent neural networks from overfitting
- ImageNet: A large-scale hierarchical image database
- (Not) Bounding the True Error
- Simplifying Neural Nets by Discovering Flat Minima
- Rethinking the Inception Architecture for Computer Vision
- Deep Residual Learning for Image Recognition
- Training Recurrent Neural Networks by Diffusion
- Reading Digits in Natural Images with Unsupervised Feature Learning
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Entropy-SGD: biasing gradient descent into wide valleys
- Understanding deep learning requires rethinking generalization
- Sharp Minima Can Generalize For Deep Nets
- Shake-Shake regularization
- Exploring Generalization in Deep Learning
- Improved Regularization of Convolutional Neural Networks with Cutout
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- MentorNet: Regularizing Very Deep Neural Networks on Corrupted Labels
- Improving Generalization Performance by Switching from Adam to SGD
Cited by
- Are All Layers Created Equal?
- Drive-specific selection in multistable mechanical networks.
- Optimize to generalize in Gaussian processes: An alternative objective based on the Rényi divergence
- Meta Pseudo Labels
- Descending through a Crowded Valley - Benchmarking Deep Learning Optimizers
- ThriftyNets : Convolutional Neural Networks with Tiny Parameter Budget
- Regularizing Neural Networks via Adversarial Model Perturbation
- Adversarial Weight Perturbation Helps Robust Generalization
- AutoDropout: Learning Dropout Patterns to Regularize Deep Networks
- Transformers in Vision: A Survey
- Design and Implementation of Deep Learning Based Contactless Authentication System Using Hand Gestures
- Multiplicative Reweighting for Robust Neural Network Optimization
- A pilot study for fragment identification using 2D NMR and deep learning
- Can Vision Transformers Learn without Natural Images?
- Facial expression and attributes recognition based on multi-task learning of lightweight neural networks
- Advanced metaheuristic optimization techniques in applications of deep neural networks: a review
- AdaBoost and robust one-bit compressed sensing
- Vision Transformers are Robust Learners
- Multi-View Fine-Grained Vehicle Classification with Multi-Loss Learning
- Understanding and Improvement of Adversarial Training for Network Embedding from an Optimization Perspective
Related papers
- Deep Residual Learning for Image Recognition
- Learning Multiple Layers of Features from Tiny Images
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- ImageNet: A large-scale hierarchical image database
- ImageNet classification with deep convolutional neural networks
- Gradient-based learning applied to document recognition
- Learning with Noisy Labels via Sparse Regularization
- GLISTER: Generalization based Data Subset Selection for Efficient and Robust Learning
- Adversarial Invariant Learning