Rethinking Attention with Performers
Explore this paper's citation graph
Summary
Performers, Transformer architectures which can estimate regular (softmax) full-rank-attention Transformers with provable accuracy, but using only linear space and time complexity, without relying on any priors such as sparsity or low-rankness are introduced.
- Type
- preprint
- Published
- 2020-09-30
- Cited by
- 2,442
- References
- 64
- Access
- Open access
- OpenAlex
- https://openalex.org/W3091156754
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:222067132
Keywords
Softmax function, Computer science, Scalability, Transformer, Quadratic equation
References
- One billion word benchmark for measuring progress in statistical language modeling
- Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books
- The conference paper
- Parallel Prefix Computation
- Random Features for Large-Scale Kernel Machines
- Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)
- Hierarchical Attention Networks for Document Classification
- Pointer Networks
- Graph Attention Networks
- The Geometry of Random Features
- Initialization matters: Orthogonal Predictive State Recurrent Neural Networks
- ListOps: A Diagnostic Dataset for Latent Tree Learning
- SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing
- The Best of Both Worlds: Combining Recent Advances in Neural Machine Translation
- Compiling machine learning programs via high-level tracing
- Transformer-XL: Language Modeling with Longer-Term Dependency
- Deep reinforcement learning with relational inductive biases
- Music Transformer: Generating Music with Long-Term Structure
- Induction of Potent Neutralizing Antibody Responses by a Designed Protein Nanoparticle Vaccine for Respiratory Syncytial Virus
- KAMA-NNs: Low-dimensional Rotation Based Neural Networks
Cited by
- Faster Neural Network Training with Approximate Tensor Operations
- An Attentive Survey of Attention Models
- CBAG: Conditional biomedical abstract generation
- Random Features for Kernel Approximation: A Survey on Algorithms, Theory, and Beyond
- Deep Learning: Our Miraculous Year 1990-1991
- Implicit Kernel Attention
- Efficient Transformers: A Survey
- Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search
- Towards a Unified Quadrature Framework for Large-Scale Kernel Machines
- Unifying Instance and Panoptic Segmentation with Dynamic Rank-1 Convolutions
- Efficient End-to-End Speech Recognition Using Performers in Conformers
- End-to-End Object Detection with Adaptive Clustering Transformer
- Long Range Arena: A Benchmark for Efficient Transformers
- A Survey of Deep Learning Approaches for OCR and Document Understanding
- Classification by Attention: Scene Graph Classification with Prior Knowledge
- PlückerNet: Learn to Register 3D Line Reconstructions¨
- Noise-Robust End-to-End Quantum Control using Deep Autoregressive Policy Networks
- The generative capacity of probabilistic protein sequence models
- Sub-Linear Memory: How to Make Performers SLiM
- Transformers in Vision: A Survey
Related papers
- Linformer: Self-Attention with Linear Complexity
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Generating Long Sequences with Sparse Transformers
- Longformer: The Long-Document Transformer
- Reformer: The Efficient Transformer
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- Language Models are Few-Shot Learners
- Deep Residual Learning for Image Recognition