Hindsight Experience Replay
Explore this paper's citation graph
Summary
A novel technique is presented which allows sample-efficient learning from rewards which are sparse and binary and therefore avoid the need for complicated reward engineering and may be seen as a form of implicit curriculum.
- Type
- preprint
- Published
- 2017-07-05
- Cited by
- 2,855
- References
- 47
- Access
- Open access
- OpenAlex
- https://openalex.org/W2733961795
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:3532908
Keywords
Hindsight bias, Task (project management), Computer science, Reinforcement learning, Artificial intelligence
References
- Universal Value Function Approximators
- Hierarchical Reinforcement Learning Based on Subgoal Discovery and Subpolicy Specialization
- Learning to Execute
- Structure in the Space of Value Functions
- Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping
- Optimal Ordered Problem Solver
- First Experiments with PowerPlay
- A theoretical analysis of Model-Based Interval Estimation
- Acceleration of stochastic approximation by averaging
- PowerPlay: Training an Increasingly General Problem Solver by Continually Searching for the Simplest Still Unsolvable Problem
- Near-Bayesian exploration in polynomial time
- Reinforcement learning of motor skills with policy gradients
- Horde: a scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
- Self-improving reactive agents based on reinforcement learning, planning and teaching
- Human-level control through deep reinforcement learning
- Learning and development in neural networks: the importance of starting small.
- MuJoCo: A physics engine for model-based control
- Mastering the game of Go with deep neural networks and tree search
- TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems
- Curriculum learning
Cited by
- Factored Contextual Policy Search with Bayesian optimization
- Deep Reinforcement Learning: An Overview
- The Intentional Unintentional Agent: Learning to Solve Many Continuous Control Tasks Simultaneously
- Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning
- Learning Generalized Reactive Policies using Deep Neural Networks
- Shared Learning : Enhancing Reinforcement in Q-Ensembles
- Overcoming Exploration in Reinforcement Learning with Demonstrations
- Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
- Asymmetric Actor Critic for Image-Based Robot Learning
- Sim-to-Real Transfer of Robotic Control with Dynamics Randomization
- Hindsight policy gradients
- Divide-and-Conquer Reinforcement Learning
- A Deeper Look at Experience Replay
- Hierarchical Actor-Critic
- RLlib: Abstractions for Distributed Reinforcement Learning
- ScreenerNet: Learning Self-Paced Curriculum for Deep Neural Networks
- Composable Planning with Attributes
- Zero-Shot Visual Imitation
- Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration
- Temporal Difference Models: Model-Free Deep RL for Model-Based Control
Related papers
- Universal Value Function Approximators
- End-to-end training of deep visuomotor policies
- Multi-Goal Reinforcement Learning: Challenging Robotics Environments and Request for Research
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Proximal Policy Optimization Algorithms
- Curiosity-Driven Exploration by Self-Supervised Prediction
- Mastering the game of Go with deep neural networks and tree search
- Continuous control with deep reinforcement learning