Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
Explore this paper's citation graph
Summary
This work expresses the self-attention as a linear dot-product of kernel feature maps and makes use of the associativity property of matrix products to reduce the complexity from O(N) to N, where N is the sequence length.
- Type
- article
- Published
- 2020-06-29
- Cited by
- 3,095
- References
- 43
- Access
- Open access
- OpenAlex
- https://openalex.org/W3034573343
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:220250819
Keywords
Autoregressive model, Transformer, Quadratic equation, Computer science, Computational complexity theory
References
- Hierarchical Probabilistic Neural Network Language Model
- Long Short-Term Memory
- Classes for fast maximum entropy training
- A Scalable Hierarchical Distributed Language Model
- Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)
- The Design for the Wall Street Journal-based CSR Corpus
- Self-Attentional Acoustic Models
- Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
- Generating Long Sequences with Sparse Transformers
- Adaptive Attention Span in Transformers
- Stand-Alone Self-Attention in Vision Models
- Sequence to Sequence Learning with Neural Networks
- MASS: Masked Sequence to Sequence Pre-training for Language Generation
- Large Memory Layers with Product Keys
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- PixelCNN++: Improving the PixelCNN with Discretized Logistic Mixture Likelihood and Other Modifications
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Transformer Dissection: An Unified Understanding for Transformer’s Attention via the Lens of Kernel
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- Q8BERT: Quantized 8Bit BERT
Cited by
- Deep Learning: Our Miraculous Year 1990-1991
- Implicit Kernel Attention
- Deep Encoder, Shallow Decoder: Reevaluating the Speed-Quality Tradeoff in Machine Translation
- Memory Transformer
- A Mathematical Theory of Attention
- Gated Recurrent Context: Softmax-Free Attention for Online Encoder-Decoder Speech Recognition
- Linear Attention Mechanism: An Efficient Attention for Semantic Segmentation
- TensorCoder: Dimension-Wise Attention via Tensor Representation for Natural Language Modeling
- The Jazz Transformer on the Front Line: Exploring the Shortcomings of AI-composed Music through Quantitative Measures
- Compression of Deep Learning Models for Text: A Survey
- Looking for change? Roll the Dice and demand Attention
- Multi-Attention-Network for Semantic Segmentation of High-Resolution Remote Sensing Images
- Efficient Transformers: A Survey
- Cluster-Former: Clustering-based Sparse Transformer for Question Answering
- Group Equivariant Stand-Alone Self-Attention For Vision
- Learning Hard Retrieval Cross Attention for Transformer
- Rethinking Attention with Performers
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- Urban Sound Classification : striving towards a fair comparison
- Higher Order Linear Transformer