Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Explore this paper's citation graph
Summary
This paper empirically show that on the ImageNet dataset large minibatches cause optimization difficulties, but when these are addressed the trained networks exhibit good generalization and enable training visual recognition models on internet-scale data with high efficiency.
- Type
- preprint
- Published
- 2017-06-08
- Cited by
- 4,180
- References
- 43
- Access
- Open access
- OpenAlex
- https://openalex.org/W2622263826
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:13905106
Keywords
Training (meteorology), Computer science, Artificial intelligence, Geography, Meteorology
References
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- One weird trick for parallelizing convolutional neural networks
- Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Why random reshuffling beats stochastic gradient descent
- Fully convolutional networks for semantic segmentation
- A Stochastic Approximation Method
- Conversational speech recognition
- Going deeper with convolutions
- Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation
- Interprocessor collective communication library (InterCom)
- ImageNet Large Scale Visual Recognition Challenge
- Introductory Lectures on Convex Optimization - A Basic Course
- Backpropagation Applied to Handwritten Zip Code Recognition
- Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups
- ImageNet classification with deep convolutional neural networks
- Deep Residual Learning for Image Recognition
- Visualizing and Understanding Convolutional Neural Networks
- Revisiting Distributed Synchronous SGD
- Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering
Cited by
- Semantic Understanding of Scenes Through the ADE20K Dataset
- Database Meets Deep Learning: Challenges and Opportunities
- Non-negative Autoencoder with Simplified Random Neural Network
- DCFNet: Discriminant Correlation Filters Network for Visual Tracking
- Deep relaxation: partial differential equations for optimizing deep neural networks
- Train longer, generalize better: closing the generalization gap in large batch training of neural networks
- Training Quantized Nets: A Deeper Understanding
- Parle: parallelizing stochastic gradient descent
- Training a Fully Convolutional Neural Network to Route Integrated Circuits
- Stochastic, Distributed and Federated Optimization for Machine Learning
- Effective Approaches to Batch Parallelization for Dynamic Neural Network Architectures
- A robust multi-batch L-BFGS method for machine learning*
- VSE++: Improved Visual-Semantic Embeddings
- Gradient Diversity Empowers Distributed Learning
- What does fault tolerant deep learning need from MPI?
- Eyes in the Dark: Distributed Scene Understanding for Disaster Management
- Super-Convergence: Very Fast Training of Residual Networks Using Large Learning Rates
- Scaling SGD Batch Size to 32K for ImageNet Training
- An Adaptive Sampling Scheme to Efficiently Train Fully Convolutional Networks for Semantic Segmentation
- ImageNet Training in Minutes
Related papers
- ИСПОЛЬЗОВAНИЕ ПОТЕНЦИAЛA СОЦИAЛЬНЫХ ПAРТНЕРОВ В ПОДГОТОВКЕ БУДУЩИХ ПЕДAГОГОВ
- Training Systems Concept for the Armored Family of Vehicles with Consideration of the Roles of Embedded Training and Stand-Alone Training Devices
- Using DataGrid Control to Realize DataBase of Querying in VB6.0
- Study and Two Types of Typical Usage of DataGrid Web Server Control
- PACWON: A parallelizing compiler for workstations on a network