BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
Explore this paper's citation graph
Summary
Benchmarking-IR (BEIR), a robust and heterogeneous evaluation benchmark for information retrieval, is introduced to better evaluate and understand existing retrieval systems, and contributes to accelerating progress towards better robust and generalizable systems in the future.
- Type
- preprint
- Published
- 2021-04-17
- Cited by
- 1,978
- References
- 102
- Access
- Open access
- OpenAlex
- https://openalex.org/W3153833878
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:233296016
Keywords
Computer science, Leverage (statistics), Benchmark (surveying), Benchmarking, Baseline (sea)
References
- An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competition
- A large annotated corpus for learning natural language inference
- The PageRank Citation Ranking : Bringing Order to the Web
- Authoritative sources in a hyperlinked environment
- Personalised Information Retrieval: survey and classification
- Violin plots : A box plot-density trace synergism
- An empirical study of tokenization strategies for biomedical information retrieval
- Bridging the lexical chasm: statistical approaches to answer-finding
- How does clickthrough data reflect retrieval quality?
- Improved Consistent Sampling, Weighted Minhash and L1 Sketching
- Exploring Network Structure, Dynamics, and Function using NetworkX
- Time is of the essence: improving recency ranking using Twitter data
- Multi-Factor Duplicate Question Detection in Stack Overflow
- Semantic Parsing on Freebase from Question-Answer Pairs
- CQADupStack: A Benchmark Data Set for Community Question-Answering Research
- Fairness in Information Retrieval
- A Deep Relevance Matching Model for Ad-hoc Retrieval
- MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
- What do a Million News Articles Look like?
- Billion-Scale Similarity Search with GPUs
Cited by
- Simplified Data Wrangling with ir_datasets
- A Systematic Evaluation of Transfer Learning and Pseudo-labeling with BERT-based Ranking Models
- Variomes: a high recall search engine to support the curation of genomic variants
- Domain-matched Pre-training Tasks for Dense Retrieval
- QA Dataset Explosion: A Taxonomy of NLP Resources for Question Answering and Reading Comprehension
- A Statutory Article Retrieval Dataset in French
- Challenges in Generalization in Open Domain Question Answering
- Simple Entity-Centric Questions Challenge Dense Retrievers
- A proposed conceptual framework for a representational approach to information retrieval
- DS-TOD: Efficient Domain Specialization for Task-Oriented Dialog
- Zero-Shot Dense Retrieval with Momentum Adversarial Domain Invariant Representations
- Improving Query Representations for Dense Retrieval with Pseudo Relevance Feedback
- Mr. TyDi: A Multi-lingual Benchmark for Dense Retrieval
- ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction
- GPL: Generative Pseudo Labeling for Unsupervised Domain Adaptation of Dense Retrieval
- DuReader-Retrieval: A Large-scale Chinese Benchmark for Passage Retrieval from Web Search Engine
- Improving Biomedical Information Retrieval with Neural Retrievers
- Internet-augmented language models through few-shot prompting for open-domain question answering
- LaPraDoR: Unsupervised Pretrained Dense Retriever for Zero-Shot Text Retrieval
- Augmenting Pre-trained Language Models with QA-Memory for Open-Domain Question Answering
Related papers
- Learning Robust Dense Retrieval Models from Incomplete Relevance Labels
- Are We Ready for Accurate and Unbiased Fine-Grained Vehicle Classification in Realistic Environments?
- Elucidating image-to-set prediction: An analysis of models, losses and datasets
- MoleculeNet: a benchmark for molecular machine learning
- Rethinking Evaluation in ASR: Are Our Models Robust Enough?
- Geometric Stability Classification: Datasets, Metamodels, and Adversarial Attacks
- QDataSet, quantum datasets for machine learning