Ternary Neural Networks with Fine-Grained Quantization
Explore this paper's citation graph
Summary
A novel fine-grained quantization (FGQ) method to ternarize pre-trained full precision models, while also constraining activations to 8 and 4-bits is proposed, which enables a full 8/4-bit inference pipeline, with best-reported accuracy using ternary weights on ImageNet dataset.
- Type
- preprint
- Published
- 2017-05-02
- Cited by
- 112
- References
- 25
- Access
- Open access
- OpenAlex
- https://openalex.org/W2610592929
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:10562854
Keywords
Computer science, Quantization (signal processing), Algorithm, Residual neural network, Ternary operation
References
- Improving the speed of neural networks on CPUs
- Dynamically scaled fixed point arithmetic
- ImageNet: A large-scale hierarchical image database
- ImageNet Large Scale Visual Recognition Challenge
- Caffe: Convolutional Architecture for Fast Feature Embedding
- ImageNet classification with deep convolutional neural networks
- Deep Residual Learning for Image Recognition
- Convolutional Neural Networks using Logarithmic Data Representation
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- Accelerating Deep Convolutional Networks using low-precision and sparsity
- ImageNet pre-trained models with batch normalization
- FINN: A Framework for Fast, Scalable Binarized Neural Network Inference
- Incremental Network Quantization: Towards Lossless CNNs with Low-Precision Weights
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks
- Deep Learning with Limited Numerical Precision
- Trained Ternary Quantization
- Deep Learning
- Learning both Weights and Connections for Efficient Neural Network
- 8-Bit Approximations for Parallelism in Deep Learning
Cited by
- SEP-Nets: Small and Effective Pattern Networks
- Ternary Residual Networks
- Lightweight Neural Networks
- Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference
- Mixed Precision Training of Convolutional Neural Networks using Integer Operations
- PACT: Parameterized Clipping Activation for Quantized Neural Networks
- Compressing Neural Networks using the Variational Information Bottleneck
- Deep Neural Network Compression with Single and Multiple Level Quantization
- ReBNet: Residual Binarized Neural Network
- Training a Binary Weight Object Detector by Knowledge Transfer for Autonomous Driving
- Compact and Fast Machine Learning Accelerator for IoT Devices
- Efficient Contextualized Representation: Language Model Pruning for Sequence Labeling
- SYQ: Learning Symmetric Quantization for Efficient Deep Neural Networks
- A Survey on Methods and Theories of Quantized Neural Networks
- Learning Low Precision Deep Neural Networks through Regularization
- Frequency-Domain Dynamic Pruning for Convolutional Neural Networks
- On Periodic Functions as Regularizers for Quantization of Neural Networks
- Lightening the Load with Highly Accurate Storage- and Energy-Efficient LightNNs
- Dataflow-based Joint Quantization of Weights and Activations for Deep Neural Networks
- HIGHLY EFFICIENT 8-BIT LOW PRECISION INFERENCE OF CONVOLUTIONAL NEURAL NETWORKS
Related papers
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- Deep Residual Learning for Image Recognition
- Ternary Weight Networks
- Trained Ternary Quantization
- ImageNet classification with deep convolutional neural networks
- ImageNet: A large-scale hierarchical image database
- Binarized Neural Networks: Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1