Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
Explore this paper's citation graph
Summary
This work describes and analyze an apparatus for adaptively modifying the proximal function, which significantly simplifies setting a learning rate and results in regret guarantees that are provably as good as the best proximal functions that can be chosen in hindsight.
- Type
- article
- Published
- 2011-02-01
- Cited by
- 11,495
- References
- 56
- OpenAlex
- https://openalex.org/W2998508934
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:538820
Keywords
Subgradient method, Computer science, Stochastic optimization, Mathematical optimization, Artificial intelligence
References
- Proceedings of the seventh annual conference on Computational learning theory
- Composite objective mirror descent
- Adaptive and Self-Confident On-Line Learning Algorithms
- Competing in the Dark: An Efficient Algorithm for Bandit Linear Optimization
- Inequalities: Theory of majorization and its applications: by Albert W. Marshall and Ingram Olkin
- Convex analysis and minimization algorithms
- Subgradient methods for convex minimization
- Extracting certainty from uncertainty: regret bounded by variation in costs
- Review of recent development: An O( n) algorithm for quadratic knapsack problems
- Efficient projections onto the l1-ball for learning in high dimensions
- Term-Weighting Approaches in Automatic Text Retrieval
- Improved second-order bounds for prediction with expert advice
- A Second-Order Perceptron Algorithm
- Mirror descent and nonlinear projected subgradient methods for convex optimization
- Concavity of certain maps on positive definite matrices and applications to Hadamard products
- An optimal method for stochastic composite optimization
- Utilization of the operation of space dilatation in the minimization of convex functions
- An algorithm for a singly constrained class of quadratic programs subject to upper and lower bounds
- A New Approach to Variable Metric Algorithms
- Exact Convex Confidence-Weighted Learning
Cited by
- ADADELTA: An Adaptive Learning Rate Method
- Training for Fast Sequential Prediction Using Dynamic Feature Selection
- Noise Perturbation for Supervised Speech Separation
- Fast and scalable distributed deep convolutional autoencoder for fMRI big data analytics
- Fast, distributed optimization strategies for resource allocation in networks
- Statistical analysis of stochastic gradient methods for generalized linear models
- Parallelized Deep Neural Networks for Distributed Intelligent Systems
- Adaptivity and Optimism: An Improved Exponentiated Gradient Algorithm
- Distributed Stochastic Optimization of Regularized Risk via Saddle-Point Problem
- ADAPTIVE STEP-SIZES FOR REINFORCEMENT LEARNING
- Anomaly detection for internet banking using supervised learning on high dimensional data
- Transition-based Dependency Parsing Using Recursive Neural Networks
- Using Sentence Plausibility to Learn the Semantics of Transitive Verbs
- Learning with sparsity: Structures, optimization and applications
- Predicting Linguistic Structure with Incomplete and Cross-Lingual Supervision
- Interactive Learning for Sequential Decisions and Predictions
- Adaptive Multi-Compositionality for Recursive Neural Models with Applications to Sentiment Analysis
- Automatic Classification of Communicative Functions of Definiteness
- Learning Reductions That Really Work
- Scalable Automated Model Search
Related papers
- Long Short-Term Memory
- BPR: Bayesian Personalized Ranking from Implicit Feedback
- Gradient-based learning applied to document recognition
- Matrix Factorization Techniques for Recommender Systems
- Large Scale Distributed Deep Networks
- Dropout: a simple way to prevent neural networks from overfitting
- Neural Collaborative Filtering
- Factorization Machines