Identifying Generalization Properties in Neural Networks
Explore this paper's citation graph
Summary
It is proved that model generalization ability is related to the Hessian, the higher-order "smoothness" terms characterized by the Lipschitz constant of the Hessians, and the scales of the parameters.
- Type
- preprint
- Published
- 2018-09-19
- Cited by
- 50
- References
- 36
- Access
- Open access
- OpenAlex
- https://openalex.org/W2891872827
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:52309674
Keywords
Generalization, Hessian matrix, Smoothness, Metric (unit), Lipschitz continuity
References
- Generating Sequences With Recurrent Neural Networks
- Linear Coupling: An Ultimate Unification of Gradient and Mirror Descent
- Acoustic Modeling Using Deep Belief Networks
- Cubic regularization of Newton method and its global performance
- Some PAC-Bayesian Theorems
- Large-Scale Video Classification with Convolutional Neural Networks
- PAC-Bayesian model averaging
- Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition
- (Not) Bounding the True Error
- PAC-Bayes & Margins
- PAC-Bayesian Inequalities for Martingales
- ImageNet classification with deep convolutional neural networks
- Deep Neural Networks for Acoustic Modeling in Speech Recognition
- Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Entropy-SGD: biasing gradient descent into wide valleys
- Understanding deep learning requires rethinking generalization
- Sharp Minima Can Generalize For Deep Nets
- Train longer, generalize better: closing the generalization gap in large batch training of neural networks
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Cited by
- Can VAEs Generate Novel Examples?
- Normalized Flat Minima: Exploring Scale Invariant Definition of Flat Minima for Neural Networks using PAC-Bayesian Analysis
- A Scale Invariant Flatness Measure for Deep Network Minima
- Are All Layers Created Equal?
- Some limitations of norm based generalization bounds in deep neural networks
- On the Generalization Gap in Reparameterizable Reinforcement Learning
- Robust manifold broad learning system for large-scale noisy chaotic time series prediction: A perturbation perspective
- Understanding Generalization through Visualizations
- An Empirical Study of Example Forgetting during Deep Neural Network Learning
- On the Relation Between the Sharpest Directions of DNN Loss and the SGD Step Length
- BN-invariant Sharpness Regularizes the Training Model to Better Generalization
- Attentive Multi-stage Learning for Early Risk Detection of Signs of Anorexia and Self-harm on Social Media
- A Reparameterization-Invariant Flatness Measure for Deep Neural Networks
- Bridging Mode Connectivity in Loss Landscapes and Adversarial Robustness
- Feature-Robustness, Flatness and Generalization Error for Deep Neural Networks
- Wide-minima Density Hypothesis and the Explore-Exploit Learning Rate Schedule
- The role of invariance in spectral complexity-based generalization bounds.
- Dissecting Non-Vacuous Generalization Bounds based on the Mean-Field Approximation
- Deforming the Loss Surface
- Deforming the Loss Surface to Affect the Behaviour of the Optimizer
Related papers
- On the connection between WRI and FWI: Analysis of the nonlinear term in the Hessian matrix
- Higher-accuracy schemes for approximating the Hessian from electronic structure calculations in chemical dynamics simulations.
- Adjoint based Hessian evaluation for SPN modeled optical tomography
- Improved Hessian estimation for adaptive random directions stochastic approximation
- Improved Pseudo-hessian for Frequency-domain Elastic Waveform Inversion
- Use of prismatic waves in full-waveform inversion with the exact Hessian
- A Unified Metric Method of Information Loss in Privacy Preserving Data Publishing