Best Practices for Crowd-based Evaluation of German Summarization: Comparing Crowd, Expert and Automatic Evaluation
Explore this paper's citation graph
Summary
This work proposes crowdsourcing as a fast, scalable, and cost-effective alternative to expert evaluations to assess the intrinsic and extrinsic quality of summarization by comparing crowd ratings with expert ratings and automatic metrics such as ROUGE, BLEU, or BertScore on a German summarization data set.
- Type
- article
- Published
- 2020-11-01
- Cited by
- 18
- References
- 61
- Access
- Open access
- OpenAlex
- https://openalex.org/W3101960104
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:226283958
Keywords
Automatic summarization, Computer science, Crowdsourcing, Quality (philosophy), Multi-document summarization
References
- Corpus Annotation through Crowdsourcing: Towards Best Practice Guidelines
- Evaluation Measures for Text Summarization
- Better Summarization Evaluation with Word Embeddings for ROUGE
- Toward using confidence intervals to compare correlations.
- Cheap and Fast – But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks
- Mean opinion score (MOS) revisited: methods and applications, limitations and alternatives
- An Investigation into the Validity of Some Metrics for Automatically Evaluating Natural Language Generation Systems
- Recent developments in text summarization
- How reliable are annotations via crowdsourcing: a study about inter-annotator agreement for multi-label image annotation
- Mind the Gap: Dangers of Divorcing Evaluations of Summary Content from Linguistic Quality
- Book Reviews: Evaluating Natural Language Processing Systems: An Analysis and Review
- Structuring, Aggregating, and Evaluating Crowdsourced Design Critique
- The Delphi method?
- Bleu: a Method for Automatic Evaluation of Machine Translation
- Evaluating Content Selection in Summarization: The Pyramid Method
- Comparing Person- and Process-centric Strategies for Obtaining Quality Data on Amazon Mechanical Turk
- ParaEval: Using Paraphrases to Evaluate Summaries Automatically
- Meteor Universal: Language Specific Translation Evaluation for Any Target Language
- A Comparison of Features for Automatic Readability Assessment
- Whose Vote Should Count More: Optimal Integration of Labels from Labelers of Unknown Expertise
Cited by
- The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation
- On the State of German (Abstractive) Text Summarization
- How to do human evaluation: A brief introduction to user studies in NLP
- A Critical Evaluation of Evaluations for Long-form Question Answering
- Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text
- Are Experts Needed? On Human Evaluation of Counselling Reflection Generation
- PROXYQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language Models
- G-SciEdBERT: A Contextualized LLM for Science Assessment Tasks in German
- LLMs as Research Tools: Applications and Evaluations in HCI Data Work
- LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
- Summarizing Speech: A Comprehensive Survey
- Towards Hybrid Human-Machine Workflow for Natural Language Generation
- Reliability of Human Evaluation for Text Summarization: Lessons Learned and Challenges Ahead
- LLMs as Span Annotators: A Comparative Study of LLMs and Humans
- A Critical Evaluation of Evaluations for Long-form Question Answering
- S PINNING THE G OLDEN T HREAD : B ENCHMARKING L ONG -F ORM G ENERATION IN L ANGUAGE M ODELS
- Using Pre-Trained Language Models for Abstractive DBPEDIA Summarization: A Comparative Study
Related papers
- Bias in News Summarization: Measures, Pitfalls and Corpora
- Organization of Documents for Multiple Document Summarization
- A Literature Study on Different Multi-Document Summarization Techniques
- A Literature Review of Modern Multi-Document Summarization Techniques
- Chinese multi-document summarization based on Topic Detection technology
- Weighted hierarchical archetypal analysis for multi-document summarization
- Research of Summarization Extraction in Multiple Topics Document
- A survey for Multi-Document Summarization