Estimation of Non-Normalized Statistical Models by Score Matching
Explore this paper's citation graph
Summary
While the estimation of the gradient of log-density function is, in principle, a very difficult non-parametric problem, it is proved a surprising result that gives a simple formula that simplifies to a sample average of a sum of some derivatives of the log- density given by the model.
- Type
- article
- Published
- 2005-12-01
- Cited by
- 2,180
- References
- 17
- OpenAlex
- https://openalex.org/W1505878979
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:1152227
Keywords
Mathematics, Normalization (sociology), Density estimation, Probability density function, Applied mathematics
References
- On Contrastive Divergence Learning
- The theory of statistics
- Estimating Overcomplete Independent Component Bases for Image Windows
- Markov Random Field Modeling in Image Analysis
- Efficiency of pseudolikelihood estimation for simple Gaussian fields
- Sparse coding with an overcomplete basis set: a strategy employed by V1?
- Spatial Interaction and the Statistical Analysis of Lattice Systems
- Training Products of Experts by Minimizing Contrastive Divergence
- A two-layer sparse coding model learns simple and complex cell receptive fields and topography from natural images.
- Blind separation of mixture of independent sources through a quasi-maximum likelihood approach
- Information Theory, Inference, and Learning Algorithms
- A generalized Gaussian image model for edge-preserving MAP estimation
- Topographic Independent Component Analysis
- Information theory, inference, and learning algorithms
Cited by
- Local Proper Scoring Rules
- Unsupervised Feature Learning and Deep Learning: A Review and New Perspectives
- Two Distributed-State Models For Generating High-Dimensional Time Series
- Probabilistic Models of Phase Variables for Visual Representation and Neural Dynamics
- Estimating Dependency Structures for non-Gaussian Components with Linear and Energy Correlations
- Higher Order Correlations within Cortical Layers Dominate Functional Connectivity in Microcolumns
- Learning Deep Representations : Toward a better new understanding of the deep learning paradigm. (Apprentissage de représentations profondes : vers une meilleure compréhension du paradigme d'apprentissage profond)
- Gradient-free Hamiltonian Monte Carlo with Efficient Kernel Exponential Families
- Learning generative models of mid-level structure in natural images
- High-dimensional inference of graphical models using regularized score matching
- Foundations and Advances in Deep Learning
- Decomposition Techniques for Learning Graphical Models
- Provable Tensor Methods for Learning Mixtures of Classifiers
- Adaptive Parallel Tempering for Stochastic Maximum Likelihood Learning of RBMs
- A Probabilistic Approach to the Primary Visual Cortex
- A review of mean-shift algorithms for clustering
- On the Convergence Properties of Contrastive Divergence
- On autoencoder scoring
- Herding Dynamic Weights for Partially Observed Random Field Models
- Learning Representations by Maximizing Compression
Related papers
- Proper local scoring rules
- Generative Modeling by Estimating Gradients of the Data Distribution
- Bayesian Learning via Stochastic Gradient Langevin Dynamics
- A Tutorial on Energy-Based Learning
- Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
- A Fast Learning Algorithm for Deep Belief Nets
- Training restricted Boltzmann machines using approximations to the likelihood gradient
- Training Products of Experts by Minimizing Contrastive Divergence
- Generative Adversarial Nets