The Effects of Hyperparameters on SGD Training of Neural Networks
Explore this paper's citation graph
Summary
The results of large scale experiments exploring different parameters of neural network classifiers, including learning rate, batch size, and depth, and their interactions are reported.
- Type
- preprint
- Published
- 2015-08-12
- Cited by
- 64
- References
- 4
- Access
- Open access
- OpenAlex
- https://openalex.org/W1814095264
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:16489716
Keywords
Hyperparameter, Artificial neural network, Artificial intelligence, Machine learning, Training (meteorology)
References
Cited by
- No More Pesky Learning Rate Guessing Games
- FireCaffe: Near-Linear Acceleration of Deep Neural Network Training on Compute Clusters
- Optimizing deep learning hyper-parameters through an evolutionary algorithm
- Operational calculus on programming spaces and generalized tensor networks
- Cyclical Learning Rates for Training Neural Networks
- Exploring the Design Space of Deep Convolutional Neural Networks at Large Scale
- Assessment and analysis of the applicability of recurrent neural networks to natural language understanding with a focus on the problem of coreference resolution
- Operational Calculus for Differentiable Programming
- Deep learning analytics for diagnostic support of breast cancer disease management
- Model Accuracy and Runtime Tradeoff in Distributed Deep Learning: A Systematic Study
- Paleo: A Performance Model for Deep Neural Networks
- A Resizable Mini-batch Gradient Descent based on a Multi-Armed Bandit
- A Practitioners' Guide to Transfer Learning for Text Classification using Convolutional Neural Networks
- Protein Family-Specific Models Using Deep Neural Networks and Transfer Learning Improve Virtual Screening and Highlight the Need for More Data
- Measuring the Effects of Data Parallelism on Neural Network Training
- Deep learning for supernovae detection
- Applicability of Recurrent Neural Networks to Player Data Analysis in Freemium Video Games
- A Comparative Study on Regularization Strategies for Embedding-based Neural Networks
- Visual Ensemble Analysis to Study the Influence of Hyper-parameters on Training Deep Neural Networks
- A Sensitivity Analysis of Attention-Gated Convolutional Neural Networks for Sentence Classification
Related papers
- Hyperparameter optimization to improve bug prediction accuracy
- To tune or not to tune? An Approach for Recommending Important Hyperparameters
- Stealing Hyperparameters in Machine Learning
- Bayesian Optimization Machine Learning Models for True and Fake News Classification
- Hyperparameter Tuning for Overlapped Software Defect Prediction Data sets