GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
Explore this paper's citation graph
Summary
GPipe is introduced, a pipeline parallelism library that allows scaling any network that can be expressed as a sequence of layers by pipelining different sub-sequences of layers on separate accelerators, resulting in almost linear speedup when a model is partitioned across multiple accelerators.
- Type
- preprint
- Published
- 2018-11-16
- Cited by
- 2,248
- References
- 68
- Access
- Open access
- OpenAlex
- https://openalex.org/W2901299405
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:53670168
Keywords
Computer science, Speedup, Pipeline (software), Artificial neural network, Machine translation
References
- Pipelined Back-Propagation for Context-Dependent Deep Neural Networks
- One weird trick for parallelizing convolutional neural networks
- Fully convolutional networks for semantic segmentation
- Algorithm 799: revolve: an implementation of checkpointing for the reverse or adjoint mode of computational differentiation
- A bridging model for parallel computation
- Care of aged doctors
- CNN Features Off-the-Shelf: An Astounding Baseline for Recognition
- Scaling Distributed Machine Learning with the Parameter Server
- Going deeper with convolutions
- ImageNet: A large-scale hierarchical image database
- On Model Parallelization and Scheduling Strategies for Distributed Machine Learning
- Performance analysis of a pipelined backpropagation parallel algorithm
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Large Scale Distributed Deep Networks
- Rethinking the Inception Architecture for Computer Vision
- MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
- Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning
- Training Deep Nets with Sublinear Memory Cost
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Cited by
- High Performance Computing: 6th Latin American Conference, CARLA 2019, Turrialba, Costa Rica, September 25–27, 2019, Revised Selected Papers
- Regularized Evolution for Image Classifier Architecture Search
- Efficient and Robust Parallel DNN Training through Model Parallelism on Multi-GPU Platform
- Domain Adaptive Transfer Learning with Specialist Models
- Accelerated Training for CNN Distributed Deep Learning through Automatic Resource-Aware Layer Placement
- A survey of the recent architectures of deep convolutional neural networks
- MultiGrain: a unified image embedding for classes and instances
- Decoupled Greedy Learning of CNNs
- Semantic Redundancies in Image-Classification Datasets: The 10% You Don't Need
- Using Machine Learning to Guide Cognitive Modeling: A Case Study in Moral Reasoning
- NAS-Bench-101: Towards Reproducible Neural Architecture Search
- Training on the Edge: The why and the how
- Scalable Deep Learning on Distributed Infrastructures
- sharpDARTS: Faster and More Accurate Differentiable Architecture Search
- ASAP: Architecture Search, Anneal and Prune
- Low-Memory Neural Network Training: A Technical Report
- Billion-scale semi-supervised learning for image classification
- A Survey on Neural Architecture Search
- SAR object classification implementation for embedded platforms
- A Survey of Multilingual Neural Machine Translation
Related papers
- Deep Residual Learning for Image Recognition
- ImageNet classification with deep convolutional neural networks
- One weird trick for parallelizing convolutional neural networks
- Going deeper with convolutions
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- PipeDream: generalized pipeline parallelism for DNN training
- Deep Learning