Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction

Explore this paper's citation graph

Summary

A practical algorithm, bootstrapping error accumulation reduction (BEAR), is proposed and it is demonstrated that BEAR is able to learn robustly from different off-policy distributions, including random and suboptimal demonstrations, on a range of continuous control tasks.

Type
preprint
Published
2019-06-03
Cited by
1,339
References
41
Access
Open access

Keywords

Reinforcement learning, Bootstrapping (finance), Computer science, Leverage (statistics), Backup

References

Cited by

Related papers