QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation
Explore this paper's citation graph
Summary
QT-Opt is introduced, a scalable self-supervised vision-based reinforcement learning framework that can leverage over 580k real-world grasp attempts to train a deep neural network Q-function with over 1.2M parameters to perform closed-loop, real- world grasping that generalizes to 96% grasp success on unseen objects.
- Type
- preprint
- Published
- 2018-06-27
- Cited by
- 1,757
- References
- 49
- Access
- Open access
- OpenAlex
- https://openalex.org/W2810785043
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:49470584
Keywords
GRASP, Reinforcement learning, Artificial intelligence, Computer science, Leverage (statistics)
References
- Massively Parallel Methods for Deep Reinforcement Learning
- Reinforcement learning in feedback control
- Deep learning for detecting robotic grasps
- Data-Driven Grasp Synthesis—A Survey
- TD-Gammon, a Self-Teaching Backgammon Program, Achieves Master-Level Play
- Acceleration of stochastic approximation by averaging
- Pose error robust grasping from contact wrench space metrics
- Learning force control policies for compliant manipulation
- Reinforcement learning of motor skills with policy gradients
- Human-level control through deep reinforcement learning
- Deep Reinforcement Learning with Double Q-Learning
- Large Scale Distributed Deep Networks
- ShapeNet: An Information-Rich 3D Model Repository
- Supersizing self-supervision: Learning to grasp from 50K tries and 700 robot hours
- The teaching of social and community medicine. Letting students do the work: an aspect of audiovisual aids to teaching.
- Collective robot reinforcement learning with distributed asynchronous guided policy search
- Regrasping Using Tactile Perception and Supervised Policy Learning
- Deep predictive policy training using reinforcement learning
- Learning a visuomotor controller for real world robotic grasping using simulated depth images
- Grasp Pose Detection in Point Clouds
Cited by
- Continuous-Time Fitted Value Iteration for Robust Policies
- Learning task-oriented grasping for tool manipulation from simulated self-supervision
- Zero-shot Sim-to-Real Transfer with Modular Priors
- Automata Guided Reinforcement Learning With Demonstrations
- Leveraging Contact Forces for Learning to Grasp
- Time Reversal as Self-Supervision
- Recovering Robustness in Model-Free Reinforcement Learning
- One-Shot High-Fidelity Imitation: Training Large-Scale Deep Nets with RL
- Closing the Sim-to-Real Loop: Adapting Simulation Randomization with Real World Experience
- Efficient Eligibility Traces for Deep Reinforcement Learning
- Training Frankenstein's Creature to Stack: HyperTree Architecture Search
- Grasp2Vec: Learning Object Representations from Self-Supervised Grasping
- WearableDL: Wearable Internet-of-Things and Deep Learning for Big Data Analytics - Concept, Literature, and Future
- Relative Entropy Regularized Policy Iteration
- Entropic Policy Composition with Generalized Policy Improvement and Divergence Correction
- Efficient Policy Learning for Robust Robot Grasping
- Residual Policy Learning
- Sim-To-Real via Sim-To-Sim: Data-Efficient Robotic Grasping via Randomized-To-Canonical Adaptation Networks
- Competitive Experience Replay
- GridSim: A Vehicle Kinematics Engine for Deep Neuroevolutionary Control in Autonomous Driving
Related papers
- Optimal grasping based on non-dimensionalized performance indices
- 3619 把持/非把持状態に対応した多指ハンドによる物体拘束の計画(G15-2 ロボティクス・メカトロニクス(2) ハンド・マニピュレーション,21世紀地球環境革命の機械工学:人・マイクロナノ・エネルギー・環境)
- Contact Point Identification by Active Sensing in Enveloping Grasp.
- Non-dimensionalized performance indices based optimal grasping for multi-fingered hands
- Scale-dependent grasp
- CAPGrasp: An R^3× SO(2)-Equivariant Continuous Approach-Constrained Generative Grasp Sampler
- Grasp sensing for human-computer interaction