Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction
Explore this paper's citation graph
Summary
A practical algorithm, bootstrapping error accumulation reduction (BEAR), is proposed and it is demonstrated that BEAR is able to learn robustly from different off-policy distributions, including random and suboptimal demonstrations, on a range of continuous control tasks.
- Type
- preprint
- Published
- 2019-06-03
- Cited by
- 1,339
- References
- 41
- Access
- Open access
- OpenAlex
- https://openalex.org/W2947150733
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:173990380
Keywords
Reinforcement learning, Bootstrapping (finance), Computer science, Leverage (statistics), Backup
References
- Error Bounds for Approximate Value Iteration
- Emphatic Temporal-Difference Learning
- Approximately Optimal Approximate Reinforcement Learning
- Off-Policy Temporal Difference Learning with Function Approximation
- Error Bounds for Approximate Policy Iteration
- Approximate modified policy iteration and its application to the game of Tetris
- ImageNet: A large-scale hierarchical image database
- Reinforcement Learning: An Introduction
- Fitted Q-iteration in continuous action-space MDPs
- MuJoCo: A physics engine for model-based control
- Is imitation learning the route to humanoid robots?
- Deep Residual Learning for Image Recognition
- A Kernel Two-Sample Test
- The Netflix Prize
- Deep Q-learning From Demonstrations
- Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
- BDD100K: A Diverse Driving Video Database with Scalable Annotation Tooling
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation
- Off-Policy Deep Reinforcement Learning by Bootstrapping the Covariate Shift
- Soft Q-Learning with Mutual-Information Regularization
Cited by
- Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
- Striving for Simplicity in Off-policy Deep Reinforcement Learning
- Safe Policy Improvement with an Estimated Baseline Policy
- A Framework for Data-Driven Robotics
- Quantile QT-Opt for Risk-Aware Vision-Based Robotic Grasping
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
- Benchmarking Batch Deep Reinforcement Learning Algorithms
- ZPD Teaching Strategies for Deep Reinforcement Learning from Demonstrations
- IRIS: Implicit Reinforcement without Interaction at Scale for Learning Control from Offline Robot Manipulation Data
- Behavior Regularized Offline Reinforcement Learning
- Learning To Reach Goals Without Reinforcement Learning
- Reward-Conditioned Policies
- Off-policy Bandits with Deficient Support
- Qgraph-bounded Q-learning: Stabilizing Model-Free Off-Policy Deep Reinforcement Learning
- BRPO: Batch Residual Policy Optimization
- Keep Doing What Worked: Behavioral Modelling Priors for Offline Reinforcement Learning
- DisCor: Corrective Feedback in Reinforcement Learning via Distribution Correction
- An empirical investigation of the challenges of real-world reinforcement learning
- Learning Sparse Rewarded Tasks from Sub-Optimal Demonstrations
- D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Related papers
- Exploring the relationship of reward and punishment in reinforcement learning
- Towards Direct Policy Search Reinforcement Learning for Robot Control
- Reinforcement learning-based dynamic scheduling for threat evaluation
- A reinforcement learning approach to power system stabilizer
- RESEARCH ON MARKOV GAME-BASED MULTIAGENT REINFORCEMENT LEARNING MODEL AND ALGORITHMS
- Adaptive action selection using utility-based reinforcement learning
- Q-Decomposition for Reinforcement Learning Agents
- Autonomous PEV Charging Scheduling Using Dyna-Q Reinforcement Learning
- Action Selection Methods in a Robotic Reinforcement Learning Scenario