High-Dimensional Continuous Control Using Generalized Advantage Estimation

Explore this paper's citation graph

Summary

This work addresses the large number of samples typically required and the difficulty of obtaining stable and steady improvement despite the nonstationarity of the incoming data by using value functions to substantially reduce the variance of policy gradient estimates at the cost of some bias.

Type
preprint
Published
2015-06-08
Cited by
4,794
References
31
Access
Open access

Keywords

Estimator, Artificial neural network, Reinforcement learning, Computer science, Variance (accounting)

References

Cited by

Related papers