DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
Explore this paper's citation graph
Summary
DoReFa-Net, a method to train convolutional neural networks that have low bitwidth weights and activations using low bit width parameter gradients, is proposed and can achieve comparable prediction accuracy as 32-bit counterparts.
- Type
- preprint
- Published
- 2016-06-20
- Cited by
- 2,291
- References
- 31
- Access
- Open access
- OpenAlex
- https://openalex.org/W2469490737
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:14395129
Keywords
Computer science, Convolutional neural network, Net (polyhedron), Convolution (computer science), Inference
References
- Improving the speed of neural networks on CPUs
- Training deep neural networks with low precision multiplications
- Compressing Deep Convolutional Networks using Vector Quantization
- Scaling up machine learning: parallel and distributed approaches
- NeuFlow: Dataflow vision processing system-on-a-chip
- DaDianNao: A Machine-Learning Supercomputer
- ImageNet: A large-scale hierarchical image database
- DianNao: a small-footprint high-throughput accelerator for ubiquitous machine-learning
- Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups
- ImageNet classification with deep convolutional neural networks
- Quantized Convolutional Neural Networks for Mobile Devices
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- BinaryNet: Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1
- Adding Gradient Noise Improves Learning for Very Deep Networks
- Bitwise Neural Networks
- Reading Digits in Natural Images with Unsupervised Feature Learning
- 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
- Deep neural networks are robust to weight binarization and other non-linear distortions
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks
Cited by
- Dissecting Xeon + FPGA: Why the integration of CPUs and FPGAs makes a power difference for the datacenter: Invited Paper
- Ternary neural networks for resource-efficient AI applications
- Loss-aware Binarization of Deep Networks
- Effective Quantization Methods for Recurrent Neural Networks
- Training Bit Fully Convolutional Network for Fast Semantic Segmentation
- FINN: A Framework for Fast, Scalable Binarized Neural Network Inference
- Scaling Binarized Neural Networks on Reconfigurable Logic
- Mixed Low-precision Deep Learning Inference using Dynamic Fixed Point
- Accelerating Binarized Convolutional Neural Networks with Software-Programmable FPGAs
- Deep Learning with Low Precision by Half-Wave Gaussian Quantization
- A Power-Aware Digital Multilayer Perceptron Accelerator with On-Chip Training Based on Approximate Computing
- Incremental Network Quantization: Towards Lossless CNNs with Low-Precision Weights
- Deep Convolutional Neural Network Inference with Floating-point Weights and Fixed-point Activations
- The ZipML Framework for Training Models with End-to-End Low Precision: The Cans, the Cannots, and a Little Bit of Deep Learning
- Efficient Processing of Deep Neural Networks: A Tutorial and Survey
- Learning Convolutional Networks for Content-Weighted Image Compression
- How to Train a Compact Binary Neural Network with High Accuracy?
- More is Less: A More Complicated Network with Less Inference Complexity
- Pyramid Vector Quantization for Deep Learning
- Ternary Neural Networks with Fine-Grained Quantization
Related papers
- Physical design tradeoffs for ASIC technologies
- Structured ASIC, evolution or revolution?
- デジタル/アナログ混載ASIC設計技術 (ASICアプリケ-ションガイド--ヒット商品を生み出すASIC開発設計ノウハウ ) -- (ASIC設計ノウハウの全て)
- Testing of application specific integrated circuit in LabVIEW environment
- Structured ASIC: Methodology and comparison
- An Adaptable Real-Time Object Detection for Traffic Surveillance using R-CNN over CNN with Improved Accuracy
- The Compatible Design Between FPGA and ASIC
- Design of Convolutional Neural Network Based on Reticulated Convolution Module