Cyclical Learning Rates for Training Neural Networks
Explore this paper's citation graph
- Type
- preprint
- Published
- 2015-06-03
- Cited by
- 3,003
- References
- 35
- Access
- Open access
- OpenAlex
- https://openalex.org/W2544860310
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:15247298
Keywords
Artificial neural network, Computer science, Artificial intelligence, Schedule, Monotonic function
References
- ADADELTA: An Adaptive Learning Rate Method
- Neural Networks: Tricks of the Trade
- An Empirical Evaluation of Deep Learning on Highway Driving
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- ADASECANT: Robust Adaptive Secant Method for Stochastic Gradient
- The Effects of Hyperparameters on SGD Training of Neural Networks
- Show and tell: A neural image caption generator
- Going deeper with convolutions
- Towards End-To-End Speech Recognition with Recurrent Neural Networks
- Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation
- ImageNet Large Scale Visual Recognition Challenge
- DeepFace: Closing the Gap to Human-Level Performance in Face Verification
- Adaptive stepsizes for recursive estimation with applications in approximate dynamic programming
- Caffe: Convolutional Architecture for Fast Feature Embedding
- ImageNet classification with deep convolutional neural networks
- Deep Residual Learning for Image Recognition
- Equilibrated adaptive learning rates for non-convex optimization
- SGDR: Stochastic Gradient Descent with Warm Restarts
- A method for solving the convex programming problem with convergence rate O(1/k^2)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Cited by
- Data Compression Based on Stacked RBM-AE Model for Wireless Sensor Networks
- Intelligent upgrade of waste-activated sludge dewatering process based on artificial neural network model: Core influential factor identification and non-experimental prediction of sludge dewatering performance.
- Attentive Training: A New Training Framework for Speech Enhancement
- Machine Learning Refined: Foundations, Algorithms, and Applications
- SGDR: Stochastic Gradient Descent with Warm Restarts
- Exploring loss function topology with cyclical learning rates
- Best Practices for Applying Deep Learning to Novel Applications
- Super-Convergence: Very Fast Training of Residual Networks Using Large Learning Rates
- Three Factors Influencing Minima in SGD
- Fixing Weight Decay Regularization in Adam
- A deep probabilistic framework for heterogeneous self-supervised learning of affordances
- Fine-tuned Language Models for Text Classification
- Understanding Short-Horizon Bias in Stochastic Meta-Optimization
- Localization and classification of cell nuclei in post-neoadjuvant breast cancer surgical specimen using fully convolutional networks
- Exponential decay sine wave learning rate for fast deep neural network training
- A disciplined approach to neural network hyper-parameters: Part 1 - learning rate, batch size, momentum, and weight decay
- ResNet Sparsifier: Learning Strict Identity Mappings in Deep Residual Networks
- Universal Language Model Fine-tuning for Text Classification
- Large Scale Automated Reading of Frontal and Lateral Chest X-Rays using Dual Convolutional Neural Networks
- Reliability Map Estimation for CNN-Based Camera Model Attribution
Related papers
- Deep Residual Learning for Image Recognition
- Adam: A Method for Stochastic Optimization
- ImageNet: A large-scale hierarchical image database
- ImageNet Large Scale Visual Recognition Challenge
- A Cyclical Learning Rate Method in Deep Learning Training
- Deep Learning
- Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
- No More Pesky Learning Rate Guessing Games
- An overview of gradient descent optimization algorithms