Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference
Explore this paper's citation graph
- Type
- preprint
- Published
- 2017-12-15
- Cited by
- 4,651
- References
- 34
- Access
- Open access
- OpenAlex
- https://openalex.org/W2777406049
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:39867659
Keywords
Quantization (signal processing), Inference, Computer science, Floating point, Artificial neural network
References
- Improving the speed of neural networks on CPUs
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Compressing Deep Convolutional Networks using Vector Quantization
- Going deeper with convolutions
- ImageNet: A large-scale hierarchical image database
- ImageNet classification with deep convolutional neural networks
- Compressing Neural Networks with the Hashing Trick
- Rethinking the Inception Architecture for Computer Vision
- Deep Residual Learning for Image Recognition
- SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <1MB model size
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- google,我,萨娜
- Speed/Accuracy Trade-Offs for Modern Convolutional Object Detectors
- Incremental Network Quantization: Towards Lossless CNNs with Low-Precision Weights
- Ternary Neural Networks with Fine-Grained Quantization
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Extremely Low Bit Neural Network: Squeeze the Last Bit Out with ADMM
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks
Cited by
- Universal Deep Neural Network Compression
- A Quantization-Friendly Separable Convolution for MobileNets
- DPRed: Making Typical Activation Values Matter In Deep Learning Computing
- MobileFaceNets: Efficient CNNs for Accurate Real-time Face Verification on Mobile Devices
- Low cost and power CNN/deep learning solution for automated driving
- Quantizing Convolutional Neural Networks for Low-Power High-Throughput Inference Engines
- Implementing the Generator of DCGAN on FPGA
- Quantizing deep convolutional networks for efficient inference: A whitepaper
- Performance of Neural Network Image Classification on Mobile CPU and GPU
- ChipNet: Real-Time LiDAR Processing for Drivable Region Segmentation on an FPGA
- MnasNet: Platform-Aware Neural Architecture Search for Mobile
- Training Compact Neural Networks with Binary Weights and Low Precision Activations
- FermiNets: Learning generative machines to generate efficient neural networks via generative synthesis
- Hardware-Aware Machine Learning: Modeling and Optimization
- Discovering Low-Precision Networks Close to Full-Precision Networks for Efficient Embedded Inference
- Discretely Relaxing Continuous Variables for tractable Variational Inference
- Coverage-Guided Fuzzing for Deep Neural Networks
- NICE: Noise Injection and Clamping Estimation for Neural Network Quantization
- 2018 Low-Power Image Recognition Challenge
- ACIQ: Analytical Clipping for Integer Quantization of neural networks
Related papers
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- Deep Residual Learning for Image Recognition
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Quantizing deep convolutional networks for efficient inference: A whitepaper
- ImageNet: A large-scale hierarchical image database
- Incremental Network Quantization: Towards Lossless CNNs with Low-Precision Weights