Re-evaluating the Role of Bleu in Machine Translation Research
Explore this paper's citation graph
Summary
It is shown that an improved Bleu score is neither necessary nor sufficient for achieving an actual improvement in translation quality, and two significant counterexamples to Bleu’s correlation with human judgments of quality are given.
- Type
- article
- Published
- 2006-04-01
- Cited by
- 500
- References
- 22
- Access
- Open access
- OpenAlex
- https://openalex.org/W1489525520
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:263885694
Keywords
BLEU, Machine translation, Computer science, Translation (biology), Metric (unit)
References
- Europarl: A Parallel Corpus for Statistical Machine Translation
- A Smorgasbord of Features for Statistical Machine Translation
- Correlating automated and human assessments of machine translation quality
- Linear B System Description for the 2005 NIST MT Evaluation Exercise
- Automatic evaluation of machine translation quality using n-gram co-occurrence statistics
- Precision and Recall of Machine Translation
- Bleu: a Method for Automatic Evaluation of Machine Translation
- METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments
- Extending the BLEU MT Evaluation Method with Frequency Weightings
- Word Sense Disambiguation vs. Statistical Machine Translation
- Minimum Error Rate Training in Statistical Machine Translation
- Automatic Evaluation of Summaries Using N-gram Co-occurrence Statistics
- Discriminative Training and Maximum Entropy Models for Statistical Machine Translation
- A Unified Framework For Automatic Evaluation Using 4-Gram Co-occurrence Statistics
- Holy and unholy grails
- Pharaoh: a beam search decoder for phrase-based statistical machine translation models
- Pharaoh: A Beam Search Decoder for Phrase-Based Statistical Machine Translation Models
- Automatic Evaluation of Machine Translation Quality Using N-gram Co-Occurrence Statistics
- NIST 2005 machine translation evaluation official results
- Automatic Evaluation of Translation Quality : Outline of Methodology and Report on Pilot Experiment
Cited by
- Automatic generation of natural language summaries
- -ing words in RBMT: multilingual evaluation and exploration of pre- and post-processing solutions
- Génération de phrases multilingues par apprentissage automatique de modèles de phrases. (Multilingual Natural Language Generation using sentence models learned from corpora)
- Comparing rule-based and data-driven approaches to Spanish-to-Basque machine translation
- Recycling texts: human evaluation of example-based machine translation subtitles for DVD
- TMX Markup: A Challenge When Adapting SMT to the Localisation Environment
- An Investigation into Automatic Translation of Prepositions in IT Technical Documentation from English to Chinese
- Acquisition automatique de sens pour la désambiguïsation et la sélection lexicale en traduction
- TransBooster:black box optimisation of machine translation systems
- Sensitivity of Automated MT Evaluation Metrics on Higher Quality MT Output: BLEU vs Task-Based Evaluation Methods
- Applying Automated Metrics to Speech Translation Dialogs
- Evaluating the impact of variation in automatically generated embodied object descriptions
- Reference-based vs. task-based evaluation of human language technology
- Compound Processing for Phrase-Based Statistical Machine Translation
- Investigating the effects of controlled language on the reading and comprehension of machine translated texts: A mixed-methods approach
- Text Content and Task Performance in the Evaluation of a Natural Language Generation System
- Lexical syntax for statistical machine translation
- Learning Labelled Dependencies in Machine Translation Evaluation
- Towards Heterogeneous Automatic MT Error Analysis
- Using F-structures in machine translation evaluation
Related papers
- Statistical Significance Tests for Machine Translation Evaluation
- Europarl: A Parallel Corpus for Statistical Machine Translation
- Further Meta-Evaluation of Machine Translation
- Syntactic Features for Evaluation of Machine Translation
- A Systematic Comparison of Various Statistical Alignment Models
- ROUGE: A Package for Automatic Evaluation of Summaries
- Discriminative Training and Maximum Entropy Models for Statistical Machine Translation
- Statistical Phrase-Based Translation
- A STUDY OF TRANSLATION ERROR RATE WITH TARGETED HUMAN ANNOTATION
- (Meta-) Evaluation of Machine Translation