Ternary Neural Networks with Fine-Grained Quantization

Explore this paper's citation graph

Summary

A novel fine-grained quantization (FGQ) method to ternarize pre-trained full precision models, while also constraining activations to 8 and 4-bits is proposed, which enables a full 8/4-bit inference pipeline, with best-reported accuracy using ternary weights on ImageNet dataset.

Type
preprint
Published
2017-05-02
Cited by
112
References
25
Access
Open access

Keywords

Computer science, Quantization (signal processing), Algorithm, Residual neural network, Ternary operation

References

Cited by

Related papers