Scaling SGD Batch Size to 32K for ImageNet Training

Explore this paper's citation graph

Summary

Layer-wise Adaptive Rate Scaling (LARS) is proposed, a method to enable large-batch training to general networks or datasets, and it can scale the batch size to 32768 for ResNet50 and 8192 for AlexNet.

Type
preprint
Published
2017-08-13
Cited by
425
References
19
Access
Open access

Keywords

Speedup, Computer science, Scaling, Batch processing, Process (computing)

References

Cited by

Related papers