Scaling SGD Batch Size to 32K for ImageNet Training
Explore this paper's citation graph
Summary
Layer-wise Adaptive Rate Scaling (LARS) is proposed, a method to enable large-batch training to general networks or datasets, and it can scale the batch size to 32768 for ResNet50 and 8192 for AlexNet.
- Type
- preprint
- Published
- 2017-08-13
- Cited by
- 425
- References
- 19
- Access
- Open access
- OpenAlex
- https://openalex.org/W2749988060
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:5919268
Keywords
Speedup, Computer science, Scaling, Batch processing, Process (computing)
References
- One weird trick for parallelizing convolutional neural networks
- Efficient mini-batch training for stochastic optimization
- ImageNet: A large-scale hierarchical image database
- Caffe: Convolutional Architecture for Fast Feature Embedding
- ImageNet classification with deep convolutional neural networks
- A solvable connectionist model of immediate recall of ordered lists
- Large Scale Distributed Deep Networks
- FireCaffe: Near-Linear Acceleration of Deep Neural Network Training on Compute Clusters
- Deep Residual Learning for Image Recognition
- TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems
- Revisiting Distributed Synchronous SGD
- Optimization Methods for Large-Scale Machine Learning
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Feature Pyramid Networks for Object Detection
- Train longer, generalize better: closing the generalization gap in large batch training of neural networks
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Et al
- Proposal Scaling Distributed Machine Learning with System and Algorithm Co-design
- Fast R-CNN
Cited by
- Train longer, generalize better: closing the generalization gap in large batch training of neural networks
- Proportionate gradient updates with PercentDelta
- ImageNet Training in Minutes
- Slim-DP: A Light Communication Data Parallelism for DNN
- AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks
- Parallel Complexity of Forward and Backward Propagation
- ImageNet Training by CPU: AlexNet in 11 Minutes and ResNet-50 in 48 Minutes
- Distributed Deep Reinforcement Learning: Learn how to play Atari games in 21 minutes
- Integrated Model, Batch, and Domain Parallelism in Training Neural Networks
- Better Generalization by Efficient Trust Region Method
- Massively Parallel Hyperparameter Tuning
- Hessian-based Analysis of Large Batch Training and Robustness to Adversaries
- The Secret Sharer: Measuring Unintended Neural Network Memorization & Extracting Secrets
- SparCML: High-Performance Sparse Communication for Machine Learning
- GossipGraD: Scalable Deep Learning using Gossip Communication based Asynchronous Gradient Descent
- High Throughput Synchronous Distributed Stochastic Gradient Descent
- A closer look at batch size in mini-batch training of deep auto-encoders
- Group Normalization
- Training Tips for the Transformer Model
- Block Mean Approximation for Efficient Second Order Optimization
Related papers
- Large Batch Training of Convolutional Networks
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Deep Residual Learning for Image Recognition
- Large Scale Distributed Deep Networks
- ImageNet: A large-scale hierarchical image database
- Highly Scalable Deep Learning Training System with Mixed-Precision: Training ImageNet in Four Minutes
- Horovod: fast and easy distributed deep learning in TensorFlow
- Very Deep Convolutional Networks for Large-Scale Image Recognition