Why Self-Attention? A Targeted Evaluation of Neural Machine Translation Architectures
Explore this paper's citation graph
Summary
The experimental results show that: 1) self-attentional networks and CNNs do not outperform RNNs in modeling subject-verb agreement over long distances; 2)Self-att attentional networks perform distinctly better than RNN's and CNN's on word sense disambiguation.
- Type
- article
- Published
- 2018-08-27
- Cited by
- 277
- References
- 30
- Access
- Open access
- OpenAlex
- https://openalex.org/W2888539709
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:52100282
Keywords
Computer science, Machine translation, Recurrent neural network, Convolutional neural network, Artificial intelligence
References
- Parallel Data, Tools and Interfaces in OPUS
- Recurrent Continuous Translation Models
- Scale-Invariant Convolutional Neural Networks
- Neural Machine Translation of Rare Words with Subword Units
- Effective Approaches to Attention-based Neural Machine Translation
- Long Short-Term Memory
- Bleu: a Method for Automatic Evaluation of Machine Translation
- Finding Structure in Time
- Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation
- MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
- Assessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies
- How Grammatical is Character-level Neural Machine Translation? Assessing MT Quality with Contrastive Translation Pairs
- Comparative Study of CNN and RNN for Natural Language Processing
- Deep architectures for Neural Machine Translation
- The University of Edinburgh’s Neural MT Systems for WMT17
- Improving Word Sense Disambiguation in Neural Machine Translation with Sense Embeddings
- Using Deep Neural Networks to Learn Syntactic Agreement
- Sockeye: A Toolkit for Neural Machine Translation
- An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling
- Marian: Fast Neural Machine Translation in C++
Cited by
- An Analysis of Attention Mechanisms: The Case of Word Sense Disambiguation in Neural Machine Translation
- The Word Sense Disambiguation Test Suite at WMT18
- DTMT: A Novel Deep Transition Architecture for Neural Machine Translation
- Do Language Models Have Common Sense
- Assessing BERT's Syntactic Abilities
- Automatic Generation of Pattern-controlled Product Description in E-commerce
- Context in Neural Machine Translation: A Review of Models and Evaluations
- Listening between the Lines: Learning Personal Attributes from Conversations
- Model-less Active Compliance for Continuum Robots using Recurrent Neural Networks
- Augmenting Neural Machine Translation with Knowledge Graphs
- Linguistic Knowledge and Transferability of Contextual Representations
- Using Multi-Sense Vector Embeddings for Reverse Dictionaries
- BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer
- Personalized Context-aware Re-ranking for E-commerce Recommender Systems
- An Attentive Survey of Attention Models
- Effective Sentence Scoring Method using Bidirectional Language Model for Speech Recognition
- A Structural Probe for Finding Syntax in Word Representations
- CNNs found to jump around more skillfully than RNNs: Compositional Generalization in Seq2seq Convolutional Networks
- Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
- Attention Is (not) All You Need for Commonsense Reasoning
Related papers
- A Study on Performance Improvement of Recurrent Neural Networks Algorithm using Word Group Expansion Technique
- Fusion Recurrent Neural Network
- Gated Feedback Recurrent Neural Networks
- Approximating Stacked and Bidirectional Recurrent Architectures with the Delayed Recurrent Neural Network
- How to Construct Deep Recurrent Neural Networks
- English-Japanese Neural Machine Translation with Encoder-Decoder-Reconstructor
- Recurrent Stacking of Layers for Compact Neural Machine Translation Models