L2 Regularization versus Batch and Weight Normalization
Explore this paper's citation graph
Summary
It is shown that popular optimization methods such as ADAM only partially eliminate the influence of normalization on the learning rate, and this leads to a discussion on other ways to mitigate this issue.
- Type
- preprint
- Published
- 2017-06-16
- Cited by
- 354
- References
- 8
- Access
- Open access
- OpenAlex
- https://openalex.org/W2626017178
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:27305706
Keywords
Normalization (sociology), Overfitting, Regularization (linguistics), Deep neural networks, Computer science
References
- On the importance of initialization and momentum in deep learning
- Deep learning via Hessian-free optimization
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks
- Layer Normalization
- Learning Multiple Layers of Features from Tiny Images
- Adam: A Method for Stochastic Optimization
- Learning Multiple Layers of Features from Tiny Images
Cited by
- A Systematic Comparison of Deep Learning Architectures in an Autonomous Vehicle
- Regularisation of neural networks by enforcing Lipschitz continuity
- Do Better ImageNet Models Transfer Better?
- On layer-level control of DNN training and its impact on generalization
- The Impact of Regularization on Convolutional Neural Networks
- High-Performance Deep Neural Network-Based Tomato Plant Diseases and Pests Diagnosis System With Refinement Filter Bank
- Towards Understanding Regularization in Batch Normalization
- Strategies for Language Identification in Code-Mixed Low Resource Languages
- On the Margin Theory of Feedforward Neural Networks
- Gradient-Coherent Strong Regularization for Deep Neural Networks
- Second-order Optimization Method for Large Mini-batch: Training ResNet-50 on ImageNet in 35 Epochs
- Hyperparameters optimization on neural networks for bond trading
- Wide neural networks of any depth evolve as linear models under gradient descent
- Understanding Regularization in Batch Normalization
- Generating Audio Using Recurrent Neural Networks
- Large-Scale Distributed Second-Order Optimization Using Kronecker-Factored Approximate Curvature for Deep Convolutional Neural Networks
- Fully automated segmentation of left ventricular myocardium from 3D late gadolinium enhancement magnetic resonance images using a U-net convolutional neural network-based model
- Deep learning based approach for fully automated detection and segmentation of hard exudate from retinal images
- Fully-automated segmentation of optic disk from retinal images using deep learning techniques
- Efficient convNets for fast traffic sign recognition
Related papers
- Avoiding Overfitting in Deep Neural Networks for Clinical Opinions Generation from General Blood Test Results
- A Modified Tikhonov Regularization Method for Solving Ill-posed Problems ——(1)Constroution of Regularization
- A generalized regularization method for nonlinear ill-posed problems enhanced for nonlinear regularization terms
- A Stable Approach for Numerical Differentiation by Local Regularization Method with its Regularization Parameter Selection Strategies
- Model Regularization of Deep Neural Networks for Robust Clinical Opinions Generation from General Blood Test Results
- On New Algorithm for TDOA Location Based on Tikhonov Regularization Theory
- Improved iterative regularization for vibration-based damage detection