On the Variance of the Adaptive Learning Rate and Beyond
Explore this paper's citation graph
Summary
This work identifies a problem of the adaptive learning rate, suggests warmup works as a variance reduction technique, and proposes RAdam, a new variant of Adam, by introducing a term to rectify the variance of theadaptive learning rate.
- Type
- preprint
- Published
- 2019-08-08
- Cited by
- 2,297
- References
- 39
- Access
- Open access
- OpenAlex
- https://openalex.org/W2968917279
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:199528271
Keywords
Computer science, Artificial intelligence, Variance reduction, Robustness (evolution), Implementation
References
- ADADELTA: An Adaptive Learning Rate Method
- One billion word benchmark for measuring progress in statistical language modeling
- The conference paper
- Advances in optimizing recurrent networks
- ImageNet: A large-scale hierarchical image database
- Rethinking the Inception Architecture for Computer Vision
- Deep Residual Learning for Image Recognition
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Optimal Hyperparameters for Deep LSTM-Networks for Sequence Labeling Tasks
- DSCOVR: Randomized Primal-Dual Block Coordinate Algorithms for Asynchronous Distributed Optimization
- Fixing Weight Decay Regularization in Adam
- Training Tips for the Transformer Model
- Efficient Contextualized Representation: Language Model Pruning for Sequence Labeling
- Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks
- Incorporating Nesterov Momentum into Adam
- Adaptive Gradient Methods with Dynamic Bound of Learning Rate
- fairseq: A Fast, Extensible Toolkit for Sequence Modeling
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Reliability-aware Dynamic Feature Composition for Name Tagging
- On the Convergence of Adam and Beyond
Cited by
- Albumentations: fast and flexible image augmentations
- A learned embedding for efficient joint analysis of millions of mass spectra
- Neutron: An Implementation of the Transformer Translation Model and its Variants
- Tuning Fairness by Marginalizing Latent Target Labels
- DeepShift: Towards Multiplication-Less Neural Networks
- A Gram-Gauss-Newton Method Learning Overparameterized Deep Neural Networks for Regression Problems
- Training Neural Networks for and by Interpolation
- Use What You Have: Video retrieval using representations from collaborative experts
- Raw-to-End Name Entity Recognition in Social Media
- The Notorious Difficulty of Comparing Human and Machine Perception
- Inertial Single Vehicle Trajectory Prediction Baselines and Applications with the NGSIM Dataset
- On Loss Functions for Supervised Monaural Time-Domain Speech Enhancement
- An Adaptive Optimizer for Measurement-Frugal Variational Algorithms
- MGBPv2: Scaling Up Multi-Grid Back-Projection Networks
- AKL-ABC: An Automatic Approximate Bayesian Computation Approach Based on Kernel Learning
- Style transfer with variational autoencoders is a promising approach to RNA-Seq data harmonization and analysis
- Transformers without Tears: Improving the Normalization of Self-Attention
- Caliban: Accurate cell tracking and lineage construction in live-cell imaging experiments with deep learning
- On Empirical Comparisons of Optimizers for Deep Learning
- On the adequacy of untuned warmup for adaptive optimization
Related papers
- Deep Residual Learning for Image Recognition
- ImageNet: A large-scale hierarchical image database
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Lookahead Optimizer: k steps forward, 1 step back
- Long Short-Term Memory
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding