A Temporal-Difference Approach to Policy Gradient Estimation

Explore this paper's citation graph

Summary

This paper proposes a new approach of reconstructing the policy gradient from the start state without requiring a particular sampling strategy, and develops the first estimator that sidesteps the distribution shift issue in a model-free way.

Type
preprint
Published
2022-02-04
Cited by
4
References
38
Access
Open access

Keywords

Estimator, Realizability, Variance (accounting), Convergence (economics), Applied mathematics

References

Cited by

Related papers