Cross-lingual Language Model Pretraining
Explore this paper's citation graph
Summary
This work proposes two methods to learn cross-lingual language models (XLMs): one unsupervised that only relies on monolingual data, and one supervised that leverages parallel data with a new cross-lingsual language model objective.
- Type
- preprint
- Published
- 2019-01-22
- Cited by
- 3,034
- References
- 52
- Access
- Open access
- OpenAlex
- https://openalex.org/W2914120296
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:58981712
Keywords
BLEU, Computer science, Machine translation, Artificial intelligence, Natural language processing
References
- Recurrent neural network based language model
- Improving Vector Space Word Representations Using Multilingual Correlation
- Parallel Data, Tools and Interfaces in OPUS
- Neural Machine Translation of Rare Words with Subword Units
- A large annotated corpus for learning natural language inference
- Long Short-Term Memory
- Moses: Open Source Toolkit for Statistical Machine Translation
- Exploiting Similarities among Languages for Machine Translation
- Optimizing Chinese Word Segmentation for Machine Translation Performance
- Multilingual Models for Compositional Distributed Semantics
- Backpropagation Through Time: What It Does and How to Do It
- Distributed Representations of Words and Phrases and their Compositionality
- Open Source Toolkit for Statistical Machine Translation: Factored Translation Models and Lattice Decoding
- Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
- Exploring the Limits of Language Modeling
- “Cloze Procedure”: A New Tool for Measuring Readability
- Massively Multilingual Word Embeddings
- Normalized Word Embedding and Orthogonal Transform for Bilingual Word Translation
- Edinburgh Neural Machine Translation Systems for WMT 16
- Bridging Nonlinearities and Stochastic Regularizers with Gaussian Error Linear Units
Cited by
- A Survey of Cross-lingual Word Embedding Models
- A Survey of the Usages of Deep Learning for Natural Language Processing
- Cross-lingual Transfer Learning for Multilingual Task Oriented Dialog
- Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond
- Contextual Word Representations: A Contextual Introduction
- An Effective Approach to Unsupervised Machine Translation
- BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model
- The FLORES Evaluation Datasets for Low-Resource Machine Translation: Nepali–English and Sinhala–English
- How to (Properly) Evaluate Cross-Lingual Word Embeddings: On Strong Baselines, Comparative Analyses, and Some Misconceptions
- Polyglot Contextual Representations Improve Crosslingual Transfer
- VideoBERT: A Joint Model for Video and Language Representation Learning
- Mask-Predict: Parallel Decoding of Conditional Masked Language Models
- ERNIE: Enhanced Representation through Knowledge Integration
- Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT
- Real-time Inference in Multi-sentence Tasks with Deep Pretrained Transformers
- Measuring Semantic Abstraction of Multilingual NMT with Paraphrase Recognition and Generation Tasks
- Cross-lingual Multi-Level Adversarial Transfer to Enhance Low-Resource Name Tagging
- A Survey of Multilingual Neural Machine Translation
- Effective Cross-lingual Transfer of Neural Machine Translation Models without Shared Vocabularies
- Joint Source-Target Self Attention with Locality Constraints
Related papers
- Better Evaluation Metrics Lead to Better Machine Translation
- Statistical Machine Translation with Rule based Machine Translation
- Neural Machine Translation of Indian Languages
- Factored Statistical Machine Translation for German-English
- On Using Monolingual Corpora in Neural Machine Translation
- Improving English-to-Indian Language Neural Machine Translation Systems
- Very Deep Transformers for Neural Machine Translation
- Deep Recurrent Models with Fast-Forward Connections for Neural Machine Translation
- Robust parfda Statistical Machine Translation Results
- Human Evaluation on Statistical Machine Translation