Decision Transformer: Reinforcement Learning via Sequence Modeling
Explore this paper's citation graph
Summary
Despite its simplicity, Decision Transformer matches or exceeds the performance of state-of-the-art model-free offline RL baselines on Atari, OpenAI Gym, and Key-to-Door tasks.
- Type
- preprint
- Published
- 2021-06-02
- Cited by
- 2,470
- References
- 91
- Access
- Open access
- OpenAlex
- https://openalex.org/W3169291081
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:235294299
Keywords
Transformer, Reinforcement learning, Computer science, Scalability, Architecture
References
- Playing Atari with Deep Reinforcement Learning
- Learning from delayed rewards
- Reinforcement Learning: An Introduction
- Human-level control through deep reinforcement learning
- The Arcade Learning Environment: An Evaluation Platform for General Agents
- SeqGAN: Sequence Generative Adversarial Nets with Policy Gradient
- Controlling Linguistic Style Aspects in Neural Language Generation
- Hafez: an Interactive Poetry Generation System
- Diversity is All You Need: Learning Skills without a Reward Function
- Learning to Write with Cooperative Discriminators
- RUDDER: Return Decomposition for Delayed Rewards
- Optimizing agent behavior over long time scales by transporting value
- A Style-Based Generator Architecture for Generative Adversarial Networks
- Deep reinforcement learning with relational inductive biases
- Go-Explore: a New Approach for Hard-Exploration Problems
- Sequence Modeling of Temporal Credit Assignment for Episodic Reinforcement Learning
- Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction
- Explain Yourself! Leveraging Language Models for Commonsense Reasoning
- When to Trust Your Model: Model-Based Policy Optimization
- Dynamics-Aware Unsupervised Discovery of Skills
Cited by
- From eye-blinks to state construction: Diagnostic benchmarks for online representation learning
- DeepAltTrip: Top-K Alternative Itineraries for Trip Recommendation
- Relative molecule self-attention transformer
- SWAT: Spatial Structure Within and Among Tokens
- Look Closer: Bridging Egocentric and Third-Person Views With Transformers for Robotic Manipulation
- LISA: Learning Interpretable Skill Abstractions from Language
- Near-optimal Offline Reinforcement Learning with Linear Representation: Leveraging Variance Information with Pessimism
- Policy Architectures for Compositional Generalization in Control
- Switch Trajectory Transformer with Distributional Value Approximation for Multi-Task Reinforcement Learning
- Unsupervised Learning of Temporal Abstractions With Slot-Based Transformers
- When Should We Prefer Offline Reinforcement Learning Over Behavioral Cloning?
- Linear Complexity Randomized Self-attention Mechanism
- Experimental Standards for Deep Learning Research: A Natural Language Processing Perspective
- Can Foundation Models Perform Zero-Shot Task Specification For Robot Manipulation?
- Towards Flexible Inference in Sequential Decision Problems via Bidirectional Transformers
- Jump-Start Reinforcement Learning
- Self-Imitation Learning from Demonstrations
- Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single Image
- Diaformer: Automatic Diagnosis via Symptoms Sequence Generation
- Pessimism meets VCG: Learning Dynamic Mechanism Design via Offline Reinforcement Learning
Related papers
- Improved autoregressive model
- Auxiliary model based recursive and iterative least squares algorithm for autoregressive output error autoregressive systems
- Determining the Number of Regimes in a Threshold Autoregressive Model Using Smooth Transition Autoregressions
- Cost-Efficient Reinforcement Learning for Optimal Trade Execution on Dynamic Market Environment
- Implicit Stacked Autoregressive Model for Video Prediction
- Performance of distributed multi-agent multi-state reinforcement spectrum management using different exploration schemes