Allowing for The Grounded Use of Temporal Difference Learning in Large Ranking Models via Substate Updates
Explore this paper's citation graph
Summary
A modification of an established reinforcement learning method to facilitate the widespread use of temporal difference learning for IR: interpolated substate temporal difference (ISSTD) learning, which outperforms the current policy gradient approach for training deep neural retrieval models.
- Type
- article
- Published
- 2021-07-11
- Cited by
- 2
- References
- 42
- Access
- Open access
- OpenAlex
- https://openalex.org/W3156988171
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:235310544
Keywords
Temporal difference learning, Computer science, Reinforcement learning, Artificial intelligence, Machine learning
References
- Interpolation-based Q-learning
- Reinforcement Learning: An Introduction
- Novelty and diversity in information retrieval evaluation
- Reinforcement Learning with Function Approximation Converges to a Region
- Deep Sentence Embedding Using Long Short-Term Memory Networks: Analysis and Application to Information Retrieval
- Actor-Critic Algorithms
- Learning to Rank Answers on Large Online QA Collections
- GloVe: Global Vectors for Word Representation
- A Study of MatchPyramid Models on Ad-hoc Retrieval
- Fairness in Information Retrieval
- MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
- Task-Oriented Query Reformulation with Reinforcement Learning
- IRGAN: A Minimax Game for Unifying Generative and Discriminative Information Retrieval Models
- Adapting Markov Decision Process for Search Result Diversification
- Reinforcement Learning to Rank with Markov Decision Process
- Rainbow: Combining Improvements in Deep Reinforcement Learning
- Reinforcement Learning to Rank in E-Commerce Search Engine: Formalization, Analysis, and Application
- A Study of Reinforcement Learning for Neural Machine Translation
- Multi Page Search with Reinforcement Learning to Rank
- Successor Uncertainties: exploration and uncertainty in temporal difference learning
Cited by
Related papers
- Bi-Rank: A New Bi-Directional Ranking Method for Goods Selection
- Reinforcement learning for MDPs using temporal difference schemes
- Querywise Fair Learning to Rank through Multi-Objective Optimization
- The LambdaLoss Framework for Ranking Metric Optimization
- OrdRank: Learning to Rank with Ordered Multiple Hyperplanes
- 2D1431 Machine Learning Lab 3: Reinforcement Learning
- Adversarially Trained Environment Models Are Effective Policy Evaluators and Improvers - An Application to Information Retrieval
- Policy Learning for Fairness in Ranking
- Ranking for Relevance and Display Preferences in Complex Presentation Layouts
- Policy Learning for Fairness in Ranking