Leveraging Local and Global Patterns for Self-Attention Networks
Explore this paper's citation graph
Summary
The extensive analyses verify that the two types of contexts are complementary to each other, and the proposed hybrid attention mechanism gives highly effective improvements in their integration.
- Type
- article
- Published
- 2019-07-01
- Cited by
- 41
- References
- 31
- Access
- Open access
- OpenAlex
- https://openalex.org/W2951563833
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:196206239
Keywords
Computer science, Leverage (statistics), Sentence, Machine translation, Phrase
References
- Statistical Significance Tests for Machine Translation Evaluation
- Neural Machine Translation of Rare Words with Subword Units
- Effective Approaches to Attention-based Neural Machine Translation
- Bleu: a Method for Automatic Evaluation of Machine Translation
- Pointwise Prediction for Robust, Adaptable Japanese Morphological Analysis
- Automatic Evaluation of Translation Quality for Distant Language Pairs
- Semantic Role Labeling with Neural Network Factors
- A Decomposable Attention Model for Natural Language Inference
- Context-dependent word representation for neural machine translation
- A Context-Aware Recurrent Encoder for Neural Machine Translation
- NTT Neural Machine Translation Systems at WAT 2017
- Self-Attentional Acoustic Models
- Dissecting Contextual Word Embeddings: Architecture and Representation
- The Best of Both Worlds: Combining Recent Advances in Neural Machine Translation
- Convolutional Self-Attention Networks
- Modeling Recurrence for Transformer
- Constituency Parsing with a Self-Attentive Encoder
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Information Aggregation for Multi-Head Attention with Routing-by-Agreement
- Deep Semantic Role Labeling with Self-Attention
Cited by
- Two-Headed Monster and Crossed Co-Attention Networks
- Mixed Multi-Head Self-Attention for Neural Machine Translation
- Unsupervised Neural Dialect Translation with Commonality and Diversity Modeling
- Fixed Encoder Self-Attention Patterns in Transformer-Based Machine Translation
- Uncertainty-Aware Curriculum Learning for Neural Machine Translation
- Modeling Local and Global Contexts for Image Captioning
- OCoR: An Overlapping-Aware Code Retriever
- HeTROPY: Explainable learning diagnostics via heterogeneous maximum-entropy and multi-spatial knowledge representation
- Mimic and Conquer: Heterogeneous Tree Structure Distillation for Syntactic NLP
- Spatial-Temporal Convolutional Graph Attention Networks for Citywide Traffic Flow Forecasting
- Constraint Translation Candidates: A Bridge between Neural Query Translation and Cross-lingual Information Retrieval
- Context-Aware Cross-Attention for Non-Autoregressive Translation
- Cross Attention with Monotonic Alignment for Speech Transformer
- Self-Paced Learning for Neural Machine Translation
- MODE-LSTM: A Parameter-efficient Recurrent Network with Multi-Scale for Sentence Classification
- Investigating Self-Attention Network for Chinese Word Segmentation
- Document Graph for Neural Machine Translation
- Deformable Self-Attention for Text Classification
- Context-aware Self-Attention Networks for Natural Language Processing
- Improving BERT with Syntax-aware Local Attention
Related papers
- Children's phrase set for text input method evaluations
- Phrase-based Chinese Mongolian statistical machine translation
- Improving Neural Machine Translation through Phrase-based Forced Decoding
- A Novel Hybrid Approach to Improve Neural Machine Translation Decoding using Phrase-Based Statistical Machine Translation
- Improving Neural Machine Translation through Phrase-based Forced Decoding
- An overview of the phrase-based statistical machine translation techniques
- Interlocking Phrases in Phrase-based Statistical Machine Translation
- A Large-scale Study of Statistical Machine Translation Methods for Khmer Language
- Accuracy-Based Scoring for Phrase-Based Statistical Machine Translation