Taming Momentum in a Distributed Asynchronous Environment
Explore this paper's citation graph
Summary
DANA mitigates the gradient staleness, despite using momentum, and therefore scales to large clusters while maintaining high final accuracy and fast convergence, and is evaluated on the CIFAR and ImageNet datasets, where it outperforms existing methods.
- Type
- preprint
- Published
- 2019-07-26
- Cited by
- 26
- References
- 69
- Access
- Open access
- OpenAlex
- https://openalex.org/W2964667463
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:198953329
Keywords
Asynchronous communication, Speedup, Asynchrony (computer programming), Momentum (technical analysis), Computer science
References
- On the importance of initialization and momentum in deep learning
- Some methods of speeding up the convergence of iteration methods
- Scaling Distributed Machine Learning with the Parameter Server
- Advances in optimizing recurrent networks
- ImageNet Large Scale Visual Recognition Challenge
- Petuum: A New Platform for Distributed Machine Learning on Big Data
- Learning multiple layers of representation.
- Hogwild: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent
- Task execution time modeling for heterogeneous computing systems
- Large Scale Distributed Deep Networks
- Deep Residual Learning for Image Recognition
- Mastering the game of Go with deep neural networks and tree search
- Revisiting Distributed Synchronous SGD
- GeePS: scalable deep learning on distributed GPUs with a GPU-specialized parameter server
- Asynchrony begets momentum, with an application to deep learning
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- HogWild++: A New Mechanism for Decentralized Asynchronous Stochastic Gradient Descent
- Why Momentum Really Works
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Linearly convergent stochastic heavy ball method for minimizing generalization error
Cited by
- Gap Aware Mitigation of Gradient Staleness
- At Stability's Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks?
- ShadowSync: Performing Synchronization in the Background for Highly Scalable Distributed Training
- Pipelined Backpropagation at Scale: Training Large Models without Batches
- FedAdapt: Adaptive Offloading for IoT Devices in Federated Learning
- Gradient Compression Supercharged High-Performance Data Parallel DNN Training
- LAGA: Lagged AllReduce with Gradient Accumulation for Minimal Idle Time
- Scheduling Hyperparameters to Improve Generalization: From Centralized SGD to Asynchronous SGD
- Reducing Impacts of System Heterogeneity in Federated Learning using Weight Update Magnitudes
- ARES: Adaptive Resource-Aware Split Learning for Internet of Things
- Energy Minimization for Federated Asynchronous Learning on Battery-Powered Mobile Devices via Application Co-running
- SMEGA2: Distributed Asynchronous Deep Neural Network Training With a Single Momentum Buffer
- FLuID: Mitigating Stragglers in Federated Learning using Invariant Dropout
- Tackling Intertwined Data and Device Heterogeneities in Federated Learning with Unlimited Staleness
- A Survey on Collaborative Learning for Intelligent Autonomous Systems
- Momentum Approximation in Asynchronous Private Federated Learning
- Energy Optimization for Federated Learning on Consumer Mobile Devices With Asynchronous SGD and Application Co-Execution
- Nesterov Method for Asynchronous Pipeline Parallel Optimization
- When Device Delays Meet Data Heterogeneity in Federated AIoT Applications
- AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism
Related papers
- Faster Distributed Deep Net Training: Computation and Communication Decoupled Stochastic Gradient Descent
- Global descent replaces gradient descent to avoid local minima problem in learning with artificial neural networks
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima
- A Diffusion Theory for Deep Learning Dynamics: Stochastic Gradient Descent Escapes From Sharp Minima Exponentially Fast
- A Diffusion Theory For Minima Selection: Stochastic Gradient Descent Exponentially Favors Flat Minima
- Techniques for avoiding local minima in gradient-descent-based ID algorithms
- Accelerating Extreme Search Based on Natural Gradient Descent with Beta Distribution