Training Products of Experts by Minimizing Contrastive Divergence
Explore this paper's citation graph
Summary
A product of experts (PoE) is an interesting candidate for a perceptual system in which rapid inference is vital and generation is unnecessary because it is hard even to approximate the derivatives of the renormalization term in the combination rule.
- Type
- article
- Published
- 2002-08-01
- Cited by
- 5,667
- References
- 25
- OpenAlex
- https://openalex.org/W2116064496
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:207596505
Keywords
Latent variable, Divergence (linguistics), Inference, Computer science, Artificial intelligence
References
- Products of Hidden Markov Models
- Learning structural descriptions from examples
- Credal Networks under Maximum Entropy
- Information processing in dynamical systems: foundations of harmony theory
- The psychology of computer vision
- The "wake-sleep" algorithm for unsupervised neural networks.
- Stochastic Relaxation, Gibbs Distributions, and the Bayesian Restoration of Images
- Unsupervised learning of distributions
- Combining Probability Distributions: A Critique and an Annotated Bibliography
- Using Generative Models for Handwritten Digit Recognition
- Connectionist Learning of Belief Networks
- Bias/Variance Decompositions for Likelihood-Based Estimators
- A Maximum Entropy Approach to Natural Language Processing
- Recognizing Hand-written Digits Using Hierarchical Products of Experts
- A Gradient-Based Boosting Algorithm for Regression Problems
- Biologically Plausible Error-Driven Learning Using Local Activation Differences: The Generalized Recirculation Algorithm
- Rate-coded Restricted Boltzmann Machines for Face Recognition
- Learning Continuous Attractors in Recurrent Networks
- Learning Representations by Recirculation
- Attractor Dynamics in Feedforward Neural Networks
Cited by
- Scalability of using Restricted Boltzmann Machines for combinatorial optimization
- Epistemological Databases for Probabilistic Knowledge Base Construction
- A hierarchy of recurrent networks for speech recognition
- Complex cell pooling and the statistics of natural images
- Structure Learning in Sequential Data
- DNdisorder: predicting protein disorder using boosting and deep networks
- Value and reward based learning in neurorobots
- Learning Document Semantic Representation with Hybrid Deep Belief Network
- Event Recognition Based on Deep Learning in Chinese Texts
- Incorporating Boltzmann Machine Priors for Semantic Labeling in Images and Videos
- A Learning Framework for Winner-Take-All Networks with Stochastic Synapses
- Fast Maximum Likelihood Estimation via Equilibrium Expectation for Large Network Data
- An Adaptive Markov Random Field for Structured Compressive Sensing
- eToxPred: a machine learning-based approach to estimate the toxicity of drug candidates
- Semisupervised Deep Stacking Network with Adaptive Learning Rate Strategy for Motor Imagery EEG Recognition
- Snuba: Automating Weak Supervision to Label Training Data
- IBAL: A Probabilistic Rational Programming Language
- Image Modeling Using Tree Structured Conditional Random Fields
- Graphical Models and Overlay Networks for Reasoning about Large Distributed Systems
- Hierarchical Probabilistic Neural Network Language Model
Related papers
- PROPCNSREG: Stata module fitting a measurement model with causal indicators
- A latent-observed dissimilarity measure
- Truncated Inference for Latent Variable Optimization Problems: Application to Robust Estimation and Learning
- Variational Inference for Integrated Choice and Latent Variable (Iclv) Models: Independence from the Number of Latent Variables and Choice Options
- Variational f-divergence Minimization
- Three models for combining information from causal indicators
- Cycle-Consistent Adversarial Learning as Approximate Bayesian Inference