Stochastic Gradient Descent Performs Variational Inference, Converges to Limit Cycles for Deep Networks
Explore this paper's citation graph
- Type
- preprint
- Published
- 2017-10-30
- Cited by
- 343
- References
- 82
- Access
- Open access
- OpenAlex
- https://openalex.org/W2765393161
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:3515208
Keywords
Stochastic gradient descent, Regularization (linguistics), Applied mathematics, Mathematics, Gradient descent
References
- On derivations and solutions of master equations and asymptotic representations
- Optimal Transport: Old and New
- An Introduction to Variational Methods for Graphical Models
- Nonequilibrium steady state of a stochastic system driven by a nonlinear drift force.
- A Complete Recipe for Stochastic Gradient MCMC
- Nonlinear Fokker-Planck Equations: Fundamentals and Applications
- Thermodynamics of the general diffusion process: Equilibrium supercurrent and nonequilibrium driven circulation with dissipation
- Auto-Encoding Variational Bayes
- Potential landscape and flux framework of nonequilibrium networks: Robustness, dissipation, and coherence of biochemical oscillations
- Calculating biological behaviors of epigenetic states in the phage λ life cycle
- Reciprocal Relations in Irreversible Processes. II.
- Nonequilibrium potentials and their power-series expansions.
- Free energy and the Fokker-Planck equation
- Keeping the neural networks simple by minimizing the description length of the weights
- THE GEOMETRY OF DISSIPATIVE EVOLUTION EQUATIONS: THE POROUS MEDIUM EQUATION
- Beyond Equilibrium Thermodynamics
- On the steady-state probability distribution of nonequilibrium stochastic systems
- THE VARIATIONAL FORMULATION OF THE FOKKER-PLANCK EQUATION
- Dropout: a simple way to prevent neural networks from overfitting
- Gradient-based learning applied to document recognition
Cited by
- Deep relaxation: partial differential equations for optimizing deep neural networks
- Super-Convergence: Very Fast Training of Residual Networks Using Large Learning Rates
- A Separation Principle for Control in the Age of Deep Learning
- Three Factors Influencing Minima in SGD
- Sampling as optimization in the space of measures: The Langevin dynamics as a composite optimization problem
- The Regularization Effects of Anisotropic Noise in Stochastic Gradient Descent
- The Anisotropic Noise in Stochastic Gradient Descent: Its Behavior of Escaping from Sharp Minima and Regularization Effects
- On the Spectral Bias of Deep Neural Networks
- Emergence of Invariance and Disentanglement in Deep Representations
- TherML: Thermodynamics of Machine Learning
- A Picture of the Energy Landscape of Deep Neural Networks
- Fast, Better Training Trick - Random Gradient
- The Dynamics of Differential Learning I: Information-Dynamics and Task Reachability
- Exchangeability and Kernel Invariance in Trained MLPs
- On the Spectral Bias of Neural Networks
- Deep Frank-Wolfe For Neural Network Optimization
- Towards Theoretical Understanding of Large Batch Training in Stochastic Gradient Descent
- On the Computational Inefficiency of Large Batch Sizes for Stochastic Gradient Descent
- Deep learning for pedestrians: backpropagation in CNNs
- A Continuous-Time Analysis of Distributed Stochastic Gradient
Related papers
- Three Factors Influencing Minima in SGD
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Deep Residual Learning for Image Recognition
- Stochastic Gradient Descent as Approximate Bayesian Inference
- Opening the Black Box of Deep Neural Networks via Information
- Learning Multiple Layers of Features from Tiny Images