Revisiting Small Batch Training for Deep Neural Networks
Explore this paper's citation graph
Summary
The collected experimental results show that increasing the mini-batch size progressively reduces the range of learning rates that provide stable convergence and acceptable test performance, which contrasts with recent work advocating the use ofmini-batch sizes in the thousands.
- Type
- preprint
- Published
- 2018-04-20
- Cited by
- 774
- References
- 29
- Access
- Open access
- OpenAlex
- https://openalex.org/W2799042347
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:5032969
Keywords
Computer science, Computation, Stochastic gradient descent, Generalization, Artificial neural network
References
- On the importance of initialization and momentum in deep learning
- One weird trick for parallelizing convolutional neural networks
- Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Deep learning in neural networks: An overview
- The general inefficiency of batch training for gradient descent learning
- ImageNet classification with deep convolutional neural networks
- Large Scale Distributed Deep Networks
- Deep Residual Learning for Image Recognition
- Distributed Deep Learning Using Synchronous Stochastic Gradient Descent
- Revisiting Distributed Synchronous SGD
- Optimization Methods for Large-Scale Machine Learning
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Batch Renormalization: Towards Reducing Minibatch Dependence in Batch-Normalized Models
- Train longer, generalize better: closing the generalization gap in large batch training of neural networks
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- A Brief Survey of Deep Reinforcement Learning
- Three Factors Influencing Minima in SGD
- Group Normalization
- Efficient BackProp
Cited by
- Deep Supervised Learning Using Local Errors
- MiMatrix: A Massively Distributed Deep Learning Framework on a Petascale High-density Heterogeneous Cluster
- Combo Loss: Handling Input and Output Imbalance in Multi-Organ Segmentation
- Analysis of DAWNBench, a Time-to-Accuracy Machine Learning Performance Benchmark
- On layer-level control of DNN training and its impact on generalization
- Backdrop: Stochastic Backpropagation
- Faster SGD training by minibatch persistency
- Pushing the boundaries of parallel Deep Learning - A practical approach
- Differentially-Private "Draw and Discard" Machine Learning
- Automatic Initialization for 3D Ultrasound CT Registration During Liver Tumor Ablations
- Improving Optimization Bounds using Machine Learning: Decision Diagrams meet Deep Reinforcement Learning
- Bayesian Distributed Stochastic Gradient Descent
- Implicit Self-Regularization in Deep Neural Networks: Evidence from Random Matrix Theory and Implications for Learning
- FRAME-LEVEL PROXIMITY AND TOUCH RECOGNITION USING CAPACITIVE SENSING AND SEMI-SUPERVISED SEQUENTIAL MODELING
- A Study of Tangerine Pest Recognition Using Advanced Deep Learning Methods
- Measuring the Effects of Data Parallelism on Neural Network Training
- Kernel-Based Training of Generative Networks
- Pre-Defined Sparse Neural Networks With Hardware Acceleration
- Deep learning for pedestrians: backpropagation in CNNs
- Interpretable Predictive Modeling for Climate Variables with Weighted Lasso
Related papers
- Learning Multiple Layers of Features from Tiny Images
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Deep Residual Learning for Image Recognition
- ImageNet classification with deep convolutional neural networks
- ImageNet: A large-scale hierarchical image database
- Dropout: a simple way to prevent neural networks from overfitting
- Very Deep Convolutional Networks for Large-Scale Image Recognition