High-Dimensional Continuous Control Using Generalized Advantage Estimation
Explore this paper's citation graph
Summary
This work addresses the large number of samples typically required and the difficulty of obtaining stable and steady improvement despite the nonstationarity of the incoming data by using value functions to substantially reduce the variance of policy gradient estimates at the cost of some bias.
- Type
- preprint
- Published
- 2015-06-08
- Cited by
- 4,794
- References
- 31
- Access
- Open access
- OpenAlex
- https://openalex.org/W1191599655
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:3075448
Keywords
Estimator, Artificial neural network, Reinforcement learning, Computer science, Variance (accounting)
References
- Steps toward Artificial Intelligence
- An Analysis of Actor/Critic Algorithms Using Eligibility Traces: Reinforcement Learning with Imperfect Value Function
- Dynamic Programming and Optimal Control, Two Volume Set
- Approximate Gradient Methods in Policy-Space Optimization of Markov Reward Processes
- Neuronlike adaptive elements that can solve difficult learning control problems
- Reinforcement Learning in POMDP's via Direct Gradient Ascent
- Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping
- Reinforcement learning in feedback control
- Real-time reinforcement learning by sequential Actor-Critics and experience replay
- Principles of Behavior
- Optimal gait and form for animal locomotion
- Stochastic policy gradient reinforcement learning on a simple 3D biped
- A Natural Policy Gradient
- Bias in Natural Actor-Critic Algorithms
- Policy Gradient Methods for Reinforcement Learning with Function Approximation
- Fast biped walking with a reflexive controller and real-time policy searching
- Actor-Critic Algorithms
- MuJoCo: A physics engine for model-based control
- Natural Actor-Critic
- Boekbespreking van D.P. Bertsekas (ed.), Dynamic programming and optimal control - volume 2
Cited by
- Continuous-Time Fitted Value Iteration for Robust Policies
- Multiagent Trust Region Policy Optimization
- Dueling Network Architectures for Deep Reinforcement Learning
- Deep Attention Recurrent Q-Network
- Memory-based control with recurrent neural networks
- Easy Monotonic Policy Iteration
- HIRL: Hierarchical Inverse Reinforcement Learning for Long-Horizon Tasks with Delayed Rewards
- Review of state-of-the-arts in artificial intelligence with application to AI safety problem
- Concrete Problems in AI Safety
- Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving
- Hybrid computing using a neural network with dynamic external memory
- Learning and Transfer of Modulated Locomotor Controllers
- Deep Learning Approximation for Stochastic Control Problems
- Q-Prop: Sample-Efficient Policy Gradient with An Off-Policy Critic
- Sample Efficient Actor-Critic with Experience Replay
- Improving Policy Gradient by Exploring Under-appreciated Rewards
- Self-Critical Sequence Training for Image Captioning
- Efficient iterative policy optimization
- RL^2: Fast Reinforcement Learning via Slow Reinforcement Learning
- Expert Level Control of Ramp Metering Based on Multi-Task Deep Reinforcement Learning
Related papers
- End-to-end training of deep visuomotor policies
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Mastering the game of Go without human knowledge
- Proximal Policy Optimization Algorithms
- Emergence of Locomotion Behaviours in Rich Environments
- Benchmarking Deep Reinforcement Learning for Continuous Control
- Mastering the game of Go with deep neural networks and tree search
- Continuous control with deep reinforcement learning
- Deterministic Policy Gradient Algorithms