Why Does Unsupervised Pre-training Help Deep Learning?
Explore this paper's citation graph
Summary
The results suggest that unsupervised pre-training guides the learning towards basins of attraction of minima that support better generalization from the training data set; the evidence from these results supports a regularization explanation for the effect of pre- training.
- Type
- article
- Published
- 2010-03-01
- Cited by
- 2,516
- References
- 68
- OpenAlex
- https://openalex.org/W2997183031
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:15796526
Keywords
Artificial intelligence, Computer science, Training (meteorology), Machine learning, Deep learning
References
- Computational neuroscience : theoretical insights into brain function
- Large-scale kernel machines
- To recognize shapes, first learn to generate images.
- Sensitive periods in development : interdisciplinary perspectives
- Unsupervised Learning of Probabilistic Grammar-Markov Models for Object Categories
- Maximum mutual information estimation of hidden Markov model parameters for speech recognition
- Classification using discriminative restricted Boltzmann machines
- Overtraining, regularization and searching for a minimum, with application to neural networks
- An empirical evaluation of deep architectures on problems with many factors of variation
- A global geometric framework for nonlinear dimensionality reduction.
- Extracting and composing robust features with denoising autoencoders
- Learning Deep Architectures for AI
- Restricted Boltzmann machines for collaborative filtering
- Measuring Invariances in Deep Networks
- On the power of small-depth threshold circuits
- Sparse Feature Learning for Deep Belief Networks
- Gradient-based learning applied to document recognition
- Cluster Kernels for Semi-Supervised Learning
- Training Products of Experts by Minimizing Contrastive Divergence
- A unified architecture for natural language processing: deep neural networks with multitask learning
Cited by
- Event Recognition Based on Deep Learning in Chinese Texts
- Denoising Autoencoder Self-Organizing Map (DASOM)
- Semisupervised Text Classification by Variational Autoencoder
- Inductive bias for semi-supervised extreme learning machine
- Unsupervised Feature Learning and Deep Learning: A Review and New Perspectives
- On Random Weights and Unsupervised Feature Learning
- Parallelized Deep Neural Networks for Distributed Intelligent Systems
- Data Science for Neuroscience: The Brain as Inspiration, Model and Data Source
- An Overview of Deep-Structured Learning for Information Processing
- The student-t mixture as a natural image patch prior with application to image compression
- Quantum Machine Learning: What Quantum Computing Means to Data Mining
- Deep Machine Learning with Spatio-Temporal Inference
- Domain Adaptive Neural Networks for Object Recognition
- viSNE and Wanderlust, two algorithms for the visualization and analysis of high-dimensional single-cell data
- Deep learning via Hessian-free optimization
- Training Restricted Boltzmann Machines
- Roles of Pre-Training and Fine-Tuning in Context-Dependent DBN-HMMs for Real-World Speech Recognition
- Object tracking: Appearance modeling and feature learning
- Data and feature mixed ensemble based extreme learning machine for medical object detection and segmentation
- A fast and efficient pre-training method based on layer-by-layer maximum discrimination for deep neural networks
Related papers
- Why Does Unsupervised Pre-training Help Deep Learning?
- Deep Residual Learning for Image Recognition
- A Fast Learning Algorithm for Deep Belief Nets
- Reducing the Dimensionality of Data with Neural Networks
- Representation Learning: A Review and New Perspectives
- AN EXPLORATORY STUDY TOWARDS APPLYING AND DEMYSTIFYING DEEP LEARNING CLASSIFICATION ON BEHAVIORAL BIG DATA
- GloVe: Global Vectors for Word Representation
- A global geometric framework for nonlinear dimensionality reduction.
- Unsupervised feature learning-based encoder and adversarial networks