Just Interpolate: Kernel "Ridgeless" Regression Can Generalize
Explore this paper's citation graph
Summary
This work isolates a phenomenon of implicit regularization for minimum-norm interpolated solutions which is due to a combination of high dimensionality of the input data, curvature of the kernel function, and favorable geometric properties of the data such as an eigenvalue decay of the empirical covariance and kernel matrices.
- Type
- article
- Published
- 2018-08-01
- Cited by
- 382
- References
- 34
- Access
- Open access
- OpenAlex
- https://openalex.org/W2886836477
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:51893757
Keywords
Kernel (algebra), Curse of dimensionality, Kernel method, Curvature, Covariance
References
- A Distribution-Free Theory of Nonparametric Regression
- Kernel Methods for Pattern Analysis
- Learning with Kernels: support vector machines, regularization, optimization, and beyond
- Generalized cross-validation as a method for choosing a good ridge parameter
- Optimal Rates for the Regularized Least-Squares Algorithm
- The spectrum of kernel random matrices
- On Early Stopping in Gradient Descent Learning
- The origins of kriging
- On the limit of the largest eigenvalue of the large dimensional sample covariance matrix
- Model Selection for Regularized Least-Squares Algorithm in Learning Theory
- Regularization Networks and Support Vector Machines
- In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning
- Kernels for Vector-Valued Functions: a Review
- Best Choices for Regularization Parameters in Learning Theory: On the Bias—Variance Problem
- Concentration Inequalities: A Nonasymptotic Theory of Independence
- Spline Models for Observational Data
- Understanding deep learning requires rethinking generalization
- Implicit Regularization in Matrix Factorization
- Training Neural Networks as Learning Data-adaptive Kernels: Provable Representation and Approximation Benefits
- Overfitting or perfect fitting? Risk bounds for classification and regression rules that interpolate
Cited by
- On the Margin Theory of Feedforward Neural Networks
- Scaling description of generalization with number of parameters in deep learning
- Consistency of Interpolation with Laplace Kernels is a High-Dimensional Phenomenon
- Training Neural Networks as Learning Data-adaptive Kernels: Provable Representation and Approximation Benefits
- Generalization Error Bounds of Gradient Descent for Learning Over-Parameterized Deep ReLU Networks
- Theory III: Dynamics and Generalization in Deep Networks
- SURPRISES IN HIGH-DIMENSIONAL RIDGELESS LEAST SQUARES INTERPOLATION
- Harmless interpolation of noisy data in regression
- A Generalization Theory of Gradient Descent for Learning Over-parameterized Deep ReLU Networks
- Painless Stochastic Gradient: Interpolation, Line-Search, and Convergence Rates
- Global Minima of DNNs: The Plenty Pantry
- Kernel Truncated Randomized Ridge Regression: Optimal Rates and Low Noise Acceleration
- Regularization Matters: Generalization and Optimization of Neural Nets v.s. their Induced Kernel
- Active Learning in the Overparameterized and Interpolating Regime
- A Gram-Gauss-Newton Method Learning Overparameterized Deep Neural Networks for Regression Problems
- On the Inductive Bias of Neural Tangent Kernels
- Understanding overfitting peaks in generalization error: Analytical risk curves for l2 and l1 penalized interpolation
- Generalization Guarantees for Neural Networks via Harnessing the Low-rank Structure of the Jacobian
- Does learning require memorization? a short tale about a long tail
- Overparameterized Nonlinear Learning: Gradient Descent Takes the Shortest Path?
Related papers
No related papers recorded.