Allowing for The Grounded Use of Temporal Difference Learning in Large Ranking Models via Substate Updates

Explore this paper's citation graph

Summary

A modification of an established reinforcement learning method to facilitate the widespread use of temporal difference learning for IR: interpolated substate temporal difference (ISSTD) learning, which outperforms the current policy gradient approach for training deep neural retrieval models.

Type
article
Published
2021-07-11
Cited by
2
References
42
Access
Open access

Keywords

Temporal difference learning, Computer science, Reinforcement learning, Artificial intelligence, Machine learning

References

Cited by

Related papers