Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding
Explore this paper's citation graph
Summary
This work introduces "deep compression", a three stage pipeline: pruning, trained quantization and Huffman coding, that work together to reduce the storage requirement of neural networks by 35x to 49x without affecting their accuracy.
- Type
- preprint
- Published
- 2015-10-01
- Cited by
- 10,382
- References
- 35
- Access
- Open access
- OpenAlex
- https://openalex.org/W2119144962
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:2134321
Keywords
Huffman coding, Computer science, Quantization (signal processing), Artificial neural network, Speedup
References
- On the Construction of Huffman Trees
- Improving the speed of neural networks on CPUs
- Memory Bounded Deep Convolutional Networks
- Phoneme probability estimation with dynamic sparsely connected artificial neural networks
- Fixed point optimization of deep convolutional neural networks for object recognition
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Compressing Deep Convolutional Networks using Vector Quantization
- The conference paper
- Fixed-point feedforward deep neural network design using weights +1, 0, and −1
- Going deeper with convolutions
- Gradient-based learning applied to document recognition
- Optimal Brain Damage
- Comparing Biases for Minimal Network Construction with Back-Propagation
- Second Order Derivatives for Network Pruning: Optimal Brain Surgeon
- Deep Fried Convnets
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Predicting Parameters in Deep Learning
- ImageNet classification with deep convolutional neural networks
- Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation
- Compressing Neural Networks with the Hashing Trick
Cited by
- Combined Edge Loss UNet for Optimized Segmentation in Total Knee Arthroplasty Preoperative Planning
- FireCaffe: Near-Linear Acceleration of Deep Neural Network Training on Compute Clusters
- ACDC: A Structured Efficient Linear Layer
- Quantized Convolutional Neural Networks for Mobile Devices
- Adjustable Bounded Rectifiers: Towards Deep Binary Representations
- Resiliency of Deep Neural Networks under Quantization
- Structured Pruning of Deep Convolutional Neural Networks
- SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <1MB model size
- EIE: Efficient Inference Engine on Compressed Deep Neural Network
- Fixed Point Quantization of Deep Convolutional Networks
- Virtualizing Deep Neural Networks for Memory-Efficient Neural Network Design
- Convolutional Neural Networks using Logarithmic Data Representation
- Blending LSTMs into CNNs
- Proteus: Exploiting Numerical Precision Variability in Deep Neural Networks
- Hardware-oriented Approximation of Convolutional Neural Networks
- Optimizing convolutional neural networks on embedded platforms with OpenCL
- Ristretto: Hardware-Oriented Approximation of Convolutional Neural Networks
- Path-Normalized Optimization of Recurrent Neural Networks with ReLU Activations
- Functional Hashing for Compressing Neural Networks
- Ternary Weight Networks
Related papers
- Learning Multiple Layers of Features from Tiny Images
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Pruning Filters for Efficient ConvNets
- Binarized Neural Networks: Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1
- EIE: Efficient Inference Engine on Compressed Deep Neural Network
- SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <1MB model size
- Deep Residual Learning for Image Recognition