Neural Machine Translation of Rare Words with Subword Units
Explore this paper's citation graph
Summary
This paper introduces a simpler and more effective approach, making the NMT model capable of open-vocabulary translation by encoding rare and unknown words as sequences of subword units, and empirically shows that subword models improve over a back-off dictionary baseline for the WMT 15 translation tasks English-German and English-Russian by 1.3 BLEU.
- Type
- preprint
- Published
- 2015-08-31
- Cited by
- 8,972
- References
- 40
- Access
- Open access
- OpenAlex
- https://openalex.org/W1816313093
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:1114678
Keywords
Computer science, Natural language processing, Machine translation, Artificial intelligence, Vocabulary
References
- ADADELTA: An Adaptive Learning Rate Method
- Word hy-phen-a-tion by com-put-er
- Character-Based PSMT for Closely Related Languages
- Modelling out-of-vocabulary words for robust speech recognition
- Character-Based Pivot Translation for Under-Resourced Languages and Domains
- Recurrent Continuous Translation Models
- On the difficulty of training recurrent neural networks
- Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation
- Effective Approaches to Attention-based Neural Machine Translation
- Character-Aware Neural Language Models
- Improving SMT quality with morpho-syntactic analysis
- Empirical Methods for Compound Splitting
- Unsupervised Morphology Rivals Supervised Morphology for Arabic MT
- On Using Very Large Target Vocabulary for Neural Machine Translation
- Machine Translation without Words through Substring Alignment
- Unsupervised Multilingual Learning for Morphological Segmentation
- Can We Translate Letters?
- Unsupervised Discovery of Morphemes
- Addressing the Rare Word Problem in Neural Machine Translation
- Integrating an Unsupervised Transliteration Model into Statistical Machine Translation
Cited by
- Natural Language Understanding with Distributed Representation
- Character-based Neural Machine Translation
- Mutual Information and Diverse Decoding Improve Neural Machine Translation
- Multi-Way, Multilingual Neural Machine Translation with a Shared Attention Mechanism
- Variable-Length Word Encodings for Neural Translation Models
- Improving Neural Machine Translation Models with Monolingual Data
- Character-based Neural Machine Translation
- A Character-level Decoder without Explicit Segmentation for Neural Machine Translation
- Achieving Open Vocabulary Neural Machine Translation with Hybrid Word-Character Models
- Noisy Parallel Approximate Decoding for Conditional Recurrent Language Model
- Variational Neural Machine Translation
- The AMU-UEDIN Submission to the WMT16 News Translation Task: Attention-based NMT Models as Feature Functions in Phrase-based SMT
- Log-linear Combinations of Monolingual and Bilingual Neural Machine Translation Models for Automatic Post-Editing
- Syntactically Guided Neural Machine Translation
- The RWTH Aachen Machine Translation Systems for IWSLT 2017
- Linguistic Input Features Improve Neural Machine Translation
- First Result on Arabic Neural Machine Translation
- Edinburgh Neural Machine Translation Systems for WMT 16
- Can neural machine translation do simultaneous translation?
- Word Representation Models for Morphologically Rich Languages in Neural Machine Translation
Related papers
- Long Short-Term Memory
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Adam: A Method for Stochastic Optimization
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Bleu: a Method for Automatic Evaluation of Machine Translation
- Moses: Open Source Toolkit for Statistical Machine Translation
- RoBERTa: A Robustly Optimized BERT Pretraining Approach