Training Tips for the Transformer Model
Explore this paper's citation graph
Summary
The experiments in neural machine translation using the recent Tensor2Tensor framework and the Transformer sequence-to-sequence model are described, confirming the general mantra “more data and larger models”.
- Type
- article
- Published
- 2018-04-01
- Cited by
- 338
- References
- 30
- Access
- Open access
- OpenAlex
- https://openalex.org/W2796108585
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:4556964
Keywords
Transformer, Computer science, Machine translation, Training set, Sentence
References
- One weird trick for parallelizing convolutional neural networks
- Neural Machine Translation of Rare Words with Subword Units
- Bleu: a Method for Automatic Evaluation of Machine Translation
- chrF: character n-gram F-score for automatic MT evaluation
- The Joy of Parallelism with CzEng 1.0
- Optimization Methods for Large-Scale Machine Learning
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Fully Character-Level Neural Machine Translation without Explicit Segmentation
- Train longer, generalize better: closing the generalization gap in large batch training of neural networks
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Scaling SGD Batch Size to 32K for ImageNet Training
- Findings of the 2017 Conference on Machine Translation (WMT17)
- Three Factors Influencing Minima in SGD
- Results of the WMT17 Metrics Shared Task
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Don't Decay the Learning Rate, Increase the Batch Size
- Overview of the IWSLT 2017 Evaluation Campaign
- Layer Normalization
- MizAR 60 for Mizar 50
Cited by
- Translating Short Segments with NMT: A Case Study in English-to-Hindi
- The University of Cambridge’s Machine Translation Systems for WMT18
- An Operation Sequence Model for Explainable Neural Machine Translation
- Trivial Transfer Learning for Low-Resource Neural Machine Translation
- Neural Machine Translation based Word Transduction Mechanisms for Low-Resource Languages
- Testsuite on Czech–English Grammatical Contrasts
- CUNI Submissions in WMT18
- EvalD Reference-Less Discourse Evaluation for WMT18
- The RWTH Aachen University Supervised Machine Translation Systems for WMT 2018
- CUNI Transformer Neural MT System for WMT18
- Findings of the 2018 Conference on Machine Translation (WMT18)
- NTT’s Neural Machine Translation Systems for WMT 2018
- DiscoFuse: A Large-Scale Dataset for Discourse-Based Sentence Fusion
- CVIT-MT Systems for WAT-2018
- Competence-based Curriculum Learning for Neural Machine Translation
- Towards End-to-end Speech-to-text Translation with Two-pass Decoding
- Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
- Making Asynchronous Stochastic Gradient Descent Work for Transformers
- Attention over Heads: A Multi-Hop Attention for Neural Machine Translation
- Classification of human white blood cells using machine learning for stain-free imaging flow cytometry
Related papers
- Mantras : words of power
- The Cult of Gāyatrī: A Study
- Living Mantra
- Is this a mantra? mantra of
- Tantric Appearances and Non-Tantric Meanings: Four Systems of the “Maṇḍala of Mantra” in the Buddhist Cakrasaṃvara Literature
- Evoking Steve Jobs’s mantra of “focus and simplicity”.
- Combining PBSMT and NMT Back-translated Data for Efficient NMT
- Training Data in Statistical Machine Translation - the More, the Better?