Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Explore this paper's citation graph

Summary

A simple, efficient intra-layer model parallel approach that enables training transformer models with billions of parameters and shows that careful attention to the placement of layer normalization in BERT-like models is critical to achieving increased performance as the model size grows.

Type
preprint
Published
2019-09-17
Cited by
3,099
References
62
Access
Open access

Keywords

Parallelism (grammar), Training (meteorology), Computer science, Parallel computing, Data parallelism

References

Cited by

Related papers