Student-t policy in reinforcement learning to acquire global optimum of robot control
Explore this paper's citation graph
Summary
The student-t policy outperforms the conventional policy in four types of simulations, two of which are difficult to learn faster without sufficient exploration and the others have the local optima.
- Type
- article
- Published
- 2019-06-15
- Cited by
- 32
- References
- 41
- OpenAlex
- https://openalex.org/W2952021385
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:189863390
Keywords
Reinforcement learning, Local optimum, Computer science, Outlier, Student's t-distribution
References
- Neuronlike adaptive elements that can solve difficult learning control problems
- Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping
- ANIMAL SEARCH STRATEGIES: A QUANTITATIVE RANDOM‐WALK ANALYSIS
- Natural Gradient Works Efficiently in Learning
- A normal approximation for the chi-square distribution
- On the information matrix of the multivariate skew-t model
- Asymptotic form of the Kullback–Leibler divergence for multivariate asymmetric heavy-tailed distributions
- V-REP: A versatile and scalable robot simulation framework
- Hierarchical Relative Entropy Policy Search
- Robust Statistical Modeling Using the t Distribution
- Simple statistical gradient-following algorithms for connectionist reinforcement learning
- Reinforcement Learning: An Introduction
- A Natural Policy Gradient
- Intrinsically Motivated Reinforcement Learning
- Student-t Processes as Alternatives to Gaussian Processes
- The development of Honda humanoid robot
- Human-level control through deep reinforcement learning
- Bias in Natural Actor-Critic Algorithms
- Deterministic Policy Gradient Algorithms
- Robust Bayesian Mixture Modelling
Cited by
- Reward-Punishment Actor-Critic Algorithm Applying to Robotic Non-grasping Manipulation
- Ensemble Reinforcement Learning-Based Supervisory Control of Hybrid Electric Vehicle for Fuel Economy Improvement
- Towards Deep Robot Learning with Optimizer applicable to Non-stationary Problems
- Reinforcement learning for quadrupedal locomotion with design of continual-hierarchical curriculum
- t-Soft Update of Target Network for Deep Reinforcement Learning
- Adaptive and Multiple Time-scale Eligibility Traces for Online Deep Reinforcement Learning
- Stochastic model predictive control for energy management of power-split plug-in hybrid electric vehicles based on reinforcement learning
- Proximal Policy Optimization with Relative Pearson Divergence*
- Bottom-up multi-agent reinforcement learning by reward shaping for cooperative-competitive tasks
- Optimization algorithm for feedback and feedforward policies towards robot control robust to sensing failures
- Mutual information matrix based on asymmetric Shannon entropy for nonlinear interactions of time series
- A Novel Sparrow Search Algorithm for the Traveling Salesman Problem
- A Fuzzy Logic Reinforcement Learning Control with Spring-Damper Device for Space Robot Capturing Satellite
- Proximal Policy Optimization with Adaptive Threshold for Symmetric Relative Density Ratio
- An Adaptive Updating Method of Target Network Based on Moment Estimates for Deep Reinforcement Learning
- L2C2: Locally Lipschitz Continuous Constraint towards Stable and Smooth Reinforcement Learning
- AdaTerm: Adaptive T-Distribution Estimated Robust Moments towards Noise-Robust Stochastic Gradient Optimizer
- Design of restricted normalizing flow towards arbitrary stochastic policy with computational efficiency
- Reward Bonuses with Gain Scheduling Inspired by Iterative Deepening Search
- AdaTerm: Adaptive T-distribution estimated robust moments for Noise-Robust stochastic gradient optimization
Related papers
- On reachability and reverse reachability analysis of communicating finite state machines
- On the reachability properties of continuous-time positive systems
- Reachability and reverse reachability analysis of CFSMs
- On the parameterized complexity of approximate counting
- On the reachability properties of continuous-time positive systems
- On the reachability and nonblocking properties for parameterized discrete event systems
- Robust density modelling using the student's t-distribution for human action recognition
- Robust multiple-model LPV approach to nonlinear process identification using mixture t distributions