Sequence Modeling of Temporal Credit Assignment for Episodic Reinforcement Learning
Explore this paper's citation graph
Summary
This work introduces a new algorithm for temporal credit assignment, which learns to decompose the episodic return back to each time-step in the trajectory using deep neural networks, and finds that expressive language models such as the Transformer can be adopted for learning the importance and the dependency of states in the trajectories, therefore providing high-quality and interpretable learned reward signals.
- Type
- preprint
- Published
- 2019-05-31
- Cited by
- 41
- References
- 42
- Access
- Open access
- OpenAlex
- https://openalex.org/W2947137205
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:173188408
Keywords
Reinforcement learning, Computer science, Artificial intelligence, Deep learning, Temporal difference learning
References
- Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models
- Temporal credit assignment in reinforcement learning
- The Cross-Entropy Method for Combinatorial and Continuous Optimization
- Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping
- Apprenticeship learning via inverse reinforcement learning
- Long Short-Term Memory
- Completely Derandomized Self-Adaptation in Evolution Strategies
- Reinforcement Learning: An Introduction
- Human-level control through deep reinforcement learning
- Reward Design via Online Gradient Ascent
- Mastering the game of Go with deep neural networks and tree search
- Unifying Count-Based Exploration and Intrinsic Motivation
- Resource Management with Deep Reinforcement Learning
- Where Do Rewards Come From
- Evolution Strategies as a Scalable Alternative to Reinforcement Learning
- Molecular de-novo design through deep reinforcement learning
- Curiosity-Driven Exploration by Self-Supervised Prediction
- Emergence of Locomotion Behaviours in Rich Environments
- Proximal Policy Optimization Algorithms
- Mastering the game of Go without human knowledge
Cited by
- Agent57: Outperforming the Atari Human Benchmark
- Decision Transformer: Reinforcement Learning via Sequence Modeling
- Generation of Game Stages With Quality and Diversity by Reinforcement Learning in Turn-Based RPG
- MultiselfGAN: A Self-Guiding Neural Architecture Search Method for Generative Adversarial Networks With Multicontrollers
- A Survey on Explainable Reinforcement Learning: Concepts, Algorithms, Challenges
- STAS: Spatial-Temporal Return Decomposition for Multi-agent Reinforcement Learning
- Interpretable Reward Redistribution in Reinforcement Learning: A Causal Approach
- Research on Reward Distribution Mechanism Based on Temporal Attention in Episodic Multi-Agent Reinforcement Learning
- Require Process Control? LSTMc is all you need!
- Improved Robot Path Planning Method Based on Deep Reinforcement Learning
- When Do Transformers Shine in RL? Decoupling Memory from Credit Assignment
- Learning Multiple Coordinated Agents under Directed Acyclic Graph Constraints
- Optimum Digital Twin Response Time for Time-Sensitive Applications
- Towards Long-delayed Sparsity: Learning a Better Transformer through Reward Redistribution
- Episodic Return Decomposition by Difference of Implicitly Assigned Sub-Trajectory Reward
- Reinforcement Learning from Bagged Reward: A Transformer-based Approach for Instance-Level Reward Redistribution
- Do Transformer World Models Give Better Policy Gradients?
- STAS: Spatial-Temporal Return Decomposition for Solving Sparse Rewards Problems in Multi-agent Reinforcement Learning
- Beyond Human Preferences: Exploring Reinforcement Learning Trajectory Evaluation and Improvement through LLMs
- Provably Efficient Offline Reinforcement Learning With Trajectory-Wise Reward
Related papers
- Reinforcement learning for MDPs using temporal difference schemes
- 2D1431 Machine Learning Lab 3: Reinforcement Learning
- Improving reinforcement learning using temporal-difference network EUROCON2009
- A New Model of Reinforcement Learning, Algorithms
- Multi-step actor-critic framework for reinforcement learning in continuous control
- Evolutionary Algorithms for Reinforcement Learning
- Reinforcement learning model, algorithms and its application
- When is Realizability Sufficient for Off-Policy Reinforcement Learning?