Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Explore this paper's citation graph
Summary
A simple, efficient intra-layer model parallel approach that enables training transformer models with billions of parameters and shows that careful attention to the placement of layer normalization in BERT-like models is critical to achieving increased performance as the model size grows.
- Type
- preprint
- Published
- 2019-09-17
- Cited by
- 3,099
- References
- 62
- Access
- Open access
- OpenAlex
- https://openalex.org/W2973727699
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:202660670
Keywords
Parallelism (grammar), Training (meteorology), Computer science, Parallel computing, Data parallelism
References
- A bridging model for parallel computation
- Long Short-Term Memory
- Scaling Distributed Machine Learning with the Parameter Server
- Dropout: a simple way to prevent neural networks from overfitting
- Communication Efficient Distributed Machine Learning with the Parameter Server
- Distributed Representations of Words and Phrases and their Compositionality
- Word Representations: A Simple and General Method for Semi-Supervised Learning
- Empirical Evaluation and Combination of Advanced Language Modeling Techniques
- GloVe: Global Vectors for Word Representation
- Training Deep Nets with Sublinear Memory Cost
- Bridging Nonlinearities and Stochastic Regularizers with Gaussian Error Linear Units
- The LAMBADA dataset: Word prediction requiring a broad discourse context
- Enriching Word Vectors with Subword Information
- context2vec: Learning Generic Context Embedding with Bidirectional LSTM
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Unsupervised Pretraining for Sequence to Sequence Learning
- Learning to Generate Reviews and Discovering Sentiment
- RACE: Large-scale ReAding Comprehension Dataset From Examinations
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Learned in Translation: Contextualized Word Vectors
Cited by
- ChatGPT makes medicine easy to swallow: an exploratory case study on simplified radiology reports
- Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- Mixed Dimension Embeddings with Application to Memory-Efficient Recommendation Systems
- ZeRO: Memory Optimization Towards Training A Trillion Parameter Models
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- Transformers: State-of-the-Art Natural Language Processing
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Emergent Properties of Finetuned Language Representation Models
- Compressive Transformers for Long-Range Sequence Modelling
- CCMatrix: Mining Billions of High-Quality Parallel Sentences on the Web
- WaLDORf: Wasteless Language-model Distillation On Reading-comprehension
- Style Example-Guided Text Generation using Generative Adversarial Transformers
- Is Attention All What You Need? - An Empirical Investigation on Convolution-Based Active Memory and Self-Attention
- Multigraph Transformer for Free-Hand Sketch Recognition
- Dual Multi-head Co-attention for Multi-choice Reading Comprehension
- Learning@home: Crowdsourced Training of Large Neural Networks using Decentralized Mixture-of-Experts
- META^2: Memory-efficient taxonomic classification and abundance estimation for metagenomics with deep learning.
- Compressing Large-Scale Transformer-Based Models: A Case Study on BERT
- Training Question Answering Models from Synthetic Data