Proximal Policy Optimization Algorithms
Explore this paper's citation graph
Summary
A new family of policy gradient methods for reinforcement learning, which alternate between sampling data through interaction with the environment, and optimizing a "surrogate" objective function using stochastic gradient ascent, are proposed.
- Type
- preprint
- Published
- 2017-07-20
- Cited by
- 30,584
- References
- 14
- Access
- Open access
- OpenAlex
- https://openalex.org/W2736601468
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:28695052
Keywords
Computer science, Optimization algorithm, Algorithm, Mathematical optimization, Mathematics
References
- Approximately Optimal Approximate Reinforcement Learning
- Learning Tetris Using the Noisy Cross-Entropy Method
- Human-level control through deep reinforcement learning
- The Arcade Learning Environment: An Evaluation Platform for General Agents
- MuJoCo: A physics engine for model-based control
- Sample Efficient Actor-Critic with Experience Replay
- Emergence of Locomotion Behaviours in Rich Environments
- Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning
- OpenAI Gym
- Benchmarking Deep Reinforcement Learning for Continuous Control
- Asynchronous Methods for Deep Reinforcement Learning
- Adam: A Method for Stochastic Optimization
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- Trust Region Policy Optimization
Cited by
- Continuous-Time Fitted Value Iteration for Robust Policies
- Multiagent Trust Region Policy Optimization
- ChatGPT makes medicine easy to swallow: an exploratory case study on simplified radiology reports
- Neural Optimizer Search with Reinforcement Learning
- Teacher–Student Curriculum Learning
- Distributionally Ambiguous Optimization for Batch Bayesian Optimization
- Learning Transferable Architectures for Scalable Image Recognition
- A Machine Learning Approach to Routing
- An Information-Theoretic Optimality Principle for Deep Reinforcement Learning
- A Brief Survey of Deep Reinforcement Learning
- Mirror Descent Search and Acceleration
- Deep Learning for Video Game Playing
- Learning Sampling Distributions for Robot Motion Planning
- Deep Reinforcement Learning that Matters
- TensorFlow Agents: Efficient Batched Reinforcement Learning in TensorFlow
- Expanding Motor Skills through Relay Neural Networks
- OptionGAN: Learning Joint Reward-Policy Options using Generative Adversarial Inverse Reinforcement Learning
- Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
- Learning Complex Swarm Behaviors by Exploiting Local Communication Protocols with Deep Reinforcement Learning
- Multi-task Learning with Gradient Guided Policy Specialization
Related papers
- ИСПОЛЬЗОВAНИЕ ПОТЕНЦИAЛA СОЦИAЛЬНЫХ ПAРТНЕРОВ В ПОДГОТОВКЕ БУДУЩИХ ПЕДAГОГОВ
- A NOVEL HYBRID ALGORITHM BASED ON CROW SEARCH ALGORITHM AND WHALE OPTIMIZATION ALGORITHM FOR HIGH-DIMENSIONAL OPTIMIZATION AND FEATURE SELECTION
- A Hybrid Algorithm Based on Invasive Weed Optimization Algorithm and Grey Wolf Optimization Algorithm
- The Harris hawks optimization algorithm, salp swarm algorithm, grasshopper optimization algorithm and dragonfly algorithm for structural design optimization of vehicle components
- A Suggestion Algorithm Instituted on Invasive Weed Optimization Algorithm and Bat Optimization Algorithm
- Improving Moth-Flame Optimization Algorithm by using Slime-Mould Algorithm
- بهینه سازی انتخاب سبد سرمایه در شرایط ریسک با الگوریتم فراابتکاری ترکیبی ژنتیک (GA) و بهینه سازی شیر (LOA)
- A Hybrid Algorithm Based on Chaos Optimization and Steepest Descent Algorithm