A Critical Evaluation of Evaluations for Long-form Question Answering
Explore this paper's citation graph
Summary
This work performs the first targeted study of the evaluation of long-form answers, covering both human and automatic evaluation practices, and presents a careful analysis of experts’ evaluation, which focuses on new aspects such as the comprehensiveness of the answer.
- Type
- preprint
- Published
- 2023-05-29
- Cited by
- 159
- References
- 69
- Access
- Open access
- OpenAlex
- https://openalex.org/W4378771494
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:258960565
Keywords
Computer science, Flexibility (engineering), Preference, Question answering, Coherence (philosophical gambling strategy)
References
- Measuring nominal scale agreement among many raters.
- Bleu: a Method for Automatic Evaluation of Machine Translation
- ROUGE: A Package for Automatic Evaluation of Summaries
- The measurement of observer agreement for categorical data.
- Non-Expert Evaluation of Summarization Systems is Risky
- Statistical methods for rates and proportions
- ELI5: Long Form Question Answering
- Texygen: A Benchmarking Platform for Text Generation Models
- How to Compare Summarizers without Target Length? Pitfalls, Solutions and Re-Examination of the Neural Summarization Literature
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Evaluating the Factual Consistency of Abstractive Text Summarization
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
- The Curious Case of Neural Text Degeneration
- BERTScore: Evaluating Text Generation with BERT
- Efficient Content-Based Sparse Attention with Routing Transformers
- Scikit-learn: Machine Learning in Python
- REALM: Retrieval-Augmented Language Model Pre-Training
- Longformer: The Long-Document Transformer
- Dense Passage Retrieval for Open-Domain Question Answering
Cited by
- Using Natural Language Explanations to Rescale Human Judgments
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
- Concise Answers to Complex Questions: Summarization of Long-form Answers
- PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations
- HAGRID: A Human-LLM Collaborative Dataset for Generative Information-Seeking with Attribution
- ExpertQA: Expert-Curated Questions and Attributed Answers
- Large Language Model Alignment: A Survey
- Human Feedback is not Gold Standard
- Interpretable Long-Form Legal Question Answering with Retrieval-Augmented Large Language Models
- Assessing Large Language Models on Climate Information
- Understanding Retrieval Augmentation for Long-Form Question Answering
- PreWoMe: Exploiting Presuppositions as Working Memory for Long Form Question Answering
- Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization
- Fully Authentic Visual Question Answering Dataset from Online Communities
- Inherent limitations of LLMs regarding spatial information
- A Framework for Exploring Player Perceptions of LLM-Generated Dialogue in Commercial Video Games
- Retrieval-based Evaluation for LLMs: A Case Study in Korean Legal QA
- Reasons to Reject? Aligning Language Models with Judgments
- CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
- PROXYQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language Models
Related papers
- Overview of Question-Answering
- A Survey on Question and Answering Systems
- An Analysis of the AskMSR Question-Answering System
- Natural Language Processing based New Approach to Design Factoid Question Answering System
- Effective Question Answering Techniques and their Evaluation Metrics
- A Simple Question Answering System
- Question Answering System Analysis Based on Machine Learning
- Information retrieval and evaluation system facing chinese question answering