Linearized two-layers neural networks in high dimension
Explore this paper's citation graph
Summary
It is proved that, if both d and N are large, the behavior of these models is instead remarkably simpler, and an equally simple bound on the generalization error of Kernel Ridge Regression is obtained.
- Type
- preprint
- Published
- 2019-04-27
- Cited by
- 274
- References
- 57
- Access
- Open access
- OpenAlex
- https://openalex.org/W2941057241
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:139103528
Keywords
Combinatorics, Degree (music), Mathematics, Dimension (graph theory), Polynomial
References
- A Distribution-Free Theory of Nonparametric Regression
- Introduction to Nonparametric Estimation
- Neural Network Learning: Theoretical Foundations
- Reproducing kernel Hilbert spaces in probability and statistics
- An Introduction to Orthogonal Polynomials
- On information plus noise kernel random matrices
- Regularization Theory and Neural Networks Architectures
- On Best Approximation by Ridge Functions
- Approximation capabilities of multilayer feedforward networks
- Optimal nonlinear approximation
- Optimal Rates for the Regularized Least-Squares Algorithm
- The spectrum of kernel random matrices
- Approximation by ridge functions and neural networks
- Projection-Based Approximation and a Duality with Kernel Methods
- Dimension-independent bounds on the degree of approximation by neural networks
- Neural Networks for Optimal Approximation of Smooth and Analytic Functions
- Approximation by superpositions of a sigmoidal function
- Random Features for Large-Scale Kernel Machines
- Approximation theory of the MLP model in neural networks
- Spherical Harmonics in p Dimensions
Cited by
- Training Neural Networks as Learning Data-adaptive Kernels: Provable Representation and Approximation Benefits
- SURPRISES IN HIGH-DIMENSIONAL RIDGELESS LEAST SQUARES INTERPOLATION
- Temporal-difference learning for nonlinear value function approximation in the lazy training regime
- On the Inductive Bias of Neural Tangent Kernels
- The Convergence Rate of Neural Networks for Learned Functions of Different Frequencies
- Training Dynamics of Deep Networks using Stochastic Gradient Descent via Neural Tangent Kernel
- Inductive Bias of Gradient Descent based Adversarial Training on Separable Data
- Generalization Guarantees for Neural Networks via Harnessing the Low-rank Structure of the Jacobian
- Approximation power of random neural networks
- Limitations of Lazy Training of Two-layers Neural Networks
- A Fine-Grained Spectral Perspective on Neural Networks
- On the Risk of Minimum-Norm Interpolants and Restricted Lower Isometry of Kernels
- On Learning Over-parameterized Neural Networks: A Functional Approximation Prospective
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural Networks
- Mean-field Behaviour of Neural Tangent Kernel for Deep Neural Networks
- Gradient Dynamics of Shallow Univariate ReLU Networks
- Towards Understanding the Spectral Bias of Deep Learning
- The Local Elasticity of Neural Networks
- Implicit Bias of Gradient Descent based Adversarial Training on Separable Data
- Mean-Field Neural ODEs via Relaxed Optimal Control
Related papers
- Gradient Descent Provably Optimizes Over-parameterized Neural Networks
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks
- Generalization of ERM in Stochastic Convex Optimization: The Dimension Strikes Back
- Learning Over-Parametrized Two-Layer ReLU Neural Networks beyond NTK
- Fundamental tradeoffs between memorization and robustness in random features and neural tangent regimes
- Optimal Reduced Model Algorithms for Data-Based State Estimation
- Asymptotic behavior of ℓp-based Laplacian regularization in semi-supervised learning
- Quadrature Points via Heat Kernel Repulsion