Neural Machine Translation by Jointly Learning to Align and Translate
Explore this paper's citation graph
Summary
It is conjecture that the use of a fixed-length vector is a bottleneck in improving the performance of this basic encoder-decoder architecture, and it is proposed to extend this by allowing a model to automatically (soft-)search for parts of a source sentence that are relevant to predicting a target word, without having to form these parts as a hard segment explicitly.
- Type
- preprint
- Published
- 2014-09-01
- Cited by
- 29,941
- References
- 33
- Access
- Open access
- OpenAlex
- https://openalex.org/W2133564696
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:11212020
Keywords
Machine translation, Computer science, Transfer-based machine translation, Example-based machine translation, Sentence
References
- ADADELTA: An Adaptive Learning Rate Method
- Theano: new features and speed improvements
- Recurrent Continuous Translation Models
- Generating Sequences With Recurrent Neural Networks
- On the difficulty of training recurrent neural networks
- Sequence Transduction with Recurrent Neural Networks
- Domain Adaptation via Pseudo In-Domain Data Selection
- Overcoming the Curse of Sentence Length for Neural Machine Translation using Automatic Segmentation
- The conference paper
- Hybrid speech recognition with Deep Bidirectional LSTM
- Long Short-Term Memory
- Continuous Space Language Models for Statistical Machine Translation
- Learning long-term dependencies with gradient descent is difficult
- Bidirectional recurrent neural networks
- Statistical Phrase-Based Translation
- Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation
- On the Properties of Neural Machine Translation: Encoder–Decoder Approaches
- Continuous Space Translation Models for Phrase-Based Statistical Machine Translation
- Fast and Robust Neural Network Joint Models for Statistical Machine Translation
- Maxout Networks
Cited by
- An attention-based effective neural model for drug-drug interactions extraction
- An attention‐based BiLSTM‐CRF approach to document‐level chemical named entity recognition
- Feature Boosting Network For 3D Pose Estimation
- Semisupervised Text Classification by Variational Autoencoder
- Uncovering Thousands of New Peptides with Sequence-Mask-Search Hybrid De Novo Peptide Sequencing Framework*
- Interpretable Predictions of Clinical Outcomes with An Attention-based Recurrent Neural Network
- A real-world test of artificial intelligence infiltration of a university examinations system: A “Turing Test” case study
- Sequence Modeling using Gated Recurrent Neural Networks
- Image Captioning with an Intermediate Attributes Layer
- WordRank: Learning Word Embeddings via Robust Ranking
- Bidirectional Recurrent Neural Networks as Generative Models
- Differential Recurrent Neural Networks for Action Recognition
- Attention-Based Models for Speech Recognition
- Building End-To-End Dialogue Systems Using Generative Hierarchical Neural Network Models
- Recognizing Functions in Binaries with Neural Networks
- Video Description Generation Incorporating Spatio-Temporal Features and a Soft-Attention Mechanism
- Not All Neural Embeddings are Born Equal
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Teaching Machines to Read and Comprehend
Related papers
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Attention is All you Need
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Hierarchical Attention Networks for Document Classification
- GloVe: Global Vectors for Word Representation
- Deep Residual Learning for Image Recognition
- ImageNet classification with deep convolutional neural networks
- Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation