Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding

Explore this paper's citation graph

Summary

This work introduces "deep compression", a three stage pipeline: pruning, trained quantization and Huffman coding, that work together to reduce the storage requirement of neural networks by 35x to 49x without affecting their accuracy.

Type
preprint
Published
2015-10-01
Cited by
10,382
References
35
Access
Open access

Keywords

Huffman coding, Computer science, Quantization (signal processing), Artificial neural network, Speedup

References

Cited by

Related papers