Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Explore this paper's citation graph

Summary

This work expresses the self-attention as a linear dot-product of kernel feature maps and makes use of the associativity property of matrix products to reduce the complexity from O(N) to N, where N is the sequence length.

Type
article
Published
2020-06-29
Cited by
3,095
References
43
Access
Open access

Keywords

Autoregressive model, Transformer, Quadratic equation, Computer science, Computational complexity theory

References

Cited by

Related papers