Distilling the Knowledge in a Neural Network
Explore this paper's citation graph
Summary
This work shows that it can significantly improve the acoustic model of a heavily used commercial system by distilling the knowledge in an ensemble of models into a single model and introduces a new type of ensemble composed of one or more full models and many specialist models which learn to distinguish fine-grained classes that the full models confuse.
- Type
- preprint
- Published
- 2015-03-09
- Cited by
- 25,840
- References
- 9
- Access
- Open access
- OpenAlex
- https://openalex.org/W1821462560
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:7200347
Keywords
Artificial neural network, Computer science, Artificial intelligence
References
- Multiple Classifier Systems
- Improving neural networks by preventing co-adaptation of feature detectors
- Dropout: a simple way to prevent neural networks from overfitting
- Adaptive Mixtures of Local Experts
- Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups
- ImageNet classification with deep convolutional neural networks
- Large Scale Distributed Deep Networks
- Model compression
- Learning small-size DNN with output-distribution-based criteria
Cited by
- Unsupervised Knowledge Transfer Using Similarity Embeddings
- Unsupervised model compression for multilayer bootstrap networks
- Fast ConvNets Using Group-Wise Brain Damage
- Distilling Word Embeddings: An Encoding Approach
- Cross Modal Distillation for Supervision Transfer
- Face Verification and Open-Set Identification for Real-Time Video Applications
- Data-free Parameter Pruning for Deep Neural Networks
- A Scale Mixture Perspective of Multiplicative Noise in Neural Networks
- Deep Learning, Dark Knowledge, and Dark Matter
- Bayesian dark knowledge
- Anticipating the future by watching unlabeled video
- Recurrent neural network training with dark knowledge transfer
- Inverting Convolutional Networks with Convolutional Networks
- Learning from LDA Using Deep Neural Networks
- An Exploration of Parameter Redundancy in Deep Networks with Circulant Projections
- Knowledge Transfer Pre-training
- DeepEar: robust smartphone audio sensing in unconstrained acoustic environments using deep learning
- Learned vs. Hand-Crafted Features for Pedestrian Gender Recognition
- Channel-Level Acceleration of Deep Face Representations
- Leveraging Human Brain Activity to Improve Object Classification
Related papers
- ИСПОЛЬЗОВAНИЕ ПОТЕНЦИAЛA СОЦИAЛЬНЫХ ПAРТНЕРОВ В ПОДГОТОВКЕ БУДУЩИХ ПЕДAГОГОВ
- Using DataGrid Control to Realize DataBase of Querying in VB6.0
- Study and Two Types of Typical Usage of DataGrid Web Server Control
- PACWON: A parallelizing compiler for workstations on a network
- Bidirectional Sort and Choosing a Row to Update or Delete by Click Any Cell in DataGrid
- OpenCL-accelerated object classification in video streams using Spatial Pooler of Hierarchical Temporal Memory
- ESKVS: efficient and secure approach for keyframes-based video summarization framework
- Flexible Application of VSFlexGrid