Why Self-Attention? A Targeted Evaluation of Neural Machine Translation Architectures

Explore this paper's citation graph

Summary

The experimental results show that: 1) self-attentional networks and CNNs do not outperform RNNs in modeling subject-verb agreement over long distances; 2)Self-att attentional networks perform distinctly better than RNN's and CNN's on word sense disambiguation.

Type
article
Published
2018-08-27
Cited by
277
References
30
Access
Open access

Keywords

Computer science, Machine translation, Recurrent neural network, Convolutional neural network, Artificial intelligence

References

Cited by

Related papers