Sequence Modeling of Temporal Credit Assignment for Episodic Reinforcement Learning

Explore this paper's citation graph

Summary

This work introduces a new algorithm for temporal credit assignment, which learns to decompose the episodic return back to each time-step in the trajectory using deep neural networks, and finds that expressive language models such as the Transformer can be adopted for learning the importance and the dependency of states in the trajectories, therefore providing high-quality and interpretable learned reward signals.

Type
preprint
Published
2019-05-31
Cited by
41
References
42
Access
Open access

Keywords

Reinforcement learning, Computer science, Artificial intelligence, Deep learning, Temporal difference learning

References

Cited by

Related papers