Proximal Policy Optimization Algorithms

Explore this paper's citation graph

Summary

A new family of policy gradient methods for reinforcement learning, which alternate between sampling data through interaction with the environment, and optimizing a "surrogate" objective function using stochastic gradient ascent, are proposed.

Type
preprint
Published
2017-07-20
Cited by
30,584
References
14
Access
Open access

Keywords

Computer science, Optimization algorithm, Algorithm, Mathematical optimization, Mathematics

References

Cited by

Related papers