Reformer: The Efficient Transformer
Explore this paper's citation graph
Summary
This work replaces dot-product attention by one that uses locality-sensitive hashing and uses reversible residual layers instead of the standard residuals, which allows storing activations only once in the training process instead of several times, making the model much more memory-efficient and much faster on long sequences.
- Type
- article
- Published
- 2020-01-13
- Cited by
- 3,064
- References
- 26
- Access
- Open access
- OpenAlex
- https://openalex.org/W2994673210
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:209315300
Keywords
Transformer, Computer science, Locality, Residual, Hash function
References
- Large-scale Simple Question Answering with Memory Networks
- Practical and Optimal LSH for Angular Distance
- The Goldilocks Principle: Reading Children's Books with Explicit Memory Representations
- One-shot Learning with Memory-Augmented Neural Networks
- Hierarchical Memory Networks
- Scaling Memory-Augmented Neural Networks with Sparse Reads and Writes
- Generating Wikipedia by Summarizing Long Sequences
- A Call for Clarity in Reporting BLEU Scores
- Scaling Neural Machine Translation
- Character-Level Language Modeling with Deeper Self-Attention
- Mesh-TensorFlow: Deep Learning for Supercomputers
- Music Transformer: Generating Music with Long-Term Structure
- Generating Long Sequences with Sparse Transformers
- Low-Memory Neural Network Training: A Technical Report
- Adaptive Attention Span in Transformers
- Stand-Alone Self-Attention in Vision Models
- Augmenting Self-attention with Persistent Memory
- Large Memory Layers with Product Keys
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Memory Networks
Cited by
- Faster Neural Network Training with Approximate Tensor Operations
- An Attentive Survey of Attention Models
- A quantum search decoder for natural language processing
- Transformers: State-of-the-Art Natural Language Processing
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Blockwise Self-Attention for Long Document Understanding
- ZiMM: a deep learning model for long term adverse events with non-clinical claims data
- Efficient Content-Based Sparse Attention with Routing Transformers
- BERT-of-Theseus: Compressing BERT by Progressive Module Replacing
- Multi-variate Probabilistic Time Series Forecasting via Conditioned Normalizing Flows
- Training Large Neural Networks with Constant Memory using a New Execution Algorithm
- Permutohedral-GCN: Graph Convolutional Networks with Global Attention
- Fixed Encoder Self-Attention Patterns in Transformer-Based Machine Translation
- Sparse Sinkhorn Attention
- Compressing Large-Scale Transformer-Based Models: A Case Study on BERT
- Transformers for Limit Order Books
- Learning Directly from Grammar Compressed Text
- Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers
- SAC: Accelerating and Structuring Self-Attention via Sparse Adaptive Connection
- Telling BERT’s Full Story: from Local Attention to Global Aggregation