Long Range Arena: A Benchmark for Efficient Transformers
Explore this paper's citation graph
Summary
A systematic and unified benchmark, LRA, specifically focused on evaluating model quality under long-context scenarios is proposed, paving the way towards better understanding this class of efficient Transformer models.
- Type
- preprint
- Published
- 2020-11-08
- Cited by
- 935
- References
- 57
- Access
- Open access
- OpenAlex
- https://openalex.org/W3103682594
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:260440449
Keywords
Benchmark (surveying), Transformer, Range (aeronautics), Computer science, Engineering
References
- Parallel and serial grouping of image elements in visual perception.
- Learning Word Vectors for Sentiment Analysis
- The ACL anthology network corpus
- Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
- The LAMBADA dataset: Word prediction requiring a broad discourse context
- A Deep Relevance Matching Model for Ad-hoc Retrieval
- A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
- Constructing Datasets for Multi-hop Reading Comprehension Across Documents
- The NarrativeQA Reading Comprehension Challenge
- Generating Wikipedia by Summarizing Long Sequences
- ListOps: A Diagnostic Dataset for Latent Tree Learning
- Universal Language Model Fine-tuning for Text Classification
- Learning long-range spatial dependencies with horizontal gated-recurrent units
- HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
- Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
- Semantic Text Matching for Long-Form Documents
- Natural Questions: A Benchmark for Question Answering Research
- Generating Long Sequences with Sparse Transformers
- Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
Cited by
- Efficient Transformers: A Survey
- Rethinking Attention with Performers
- Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search
- Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with Transformers
- Sub-Linear Memory: How to Make Performers SLiM
- Match-Ignition: Plugging PageRank into Transformer for Long-form Text Matching
- Position Information in Transformers: An Overview
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image Encoding
- A Practical Survey on Faster and Lighter Transformers
- ViViT: A Video Vision Transformer
- NLQuAD: A Non-Factoid Long Question Answering Data Set
- Long-Span Dependencies in Transformer-based Summarization Systems
- FNet: Mixing Tokens with Fourier Transforms
- A Survey on Software Defect Prediction Using Deep Learning
- Memory-efficient Transformers via Top-k Attention
- Are Pretrained Convolutions Better than Pretrained Transformers?
- Long-Span Summarization via Local Attention and Content Selection
- Nyströmformer: A Nystöm-based Algorithm for Approximating Self-Attention
- H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for Sequences
Related papers
- ИСПОЛЬЗОВAНИЕ ПОТЕНЦИAЛA СОЦИAЛЬНЫХ ПAРТНЕРОВ В ПОДГОТОВКЕ БУДУЩИХ ПЕДAГОГОВ
- Theoretical Analysis of the Benchmark for Choosing Manipulative Instruments of Monetary Policies
- Exploring disk performance benchmarks
- Solutions to the Third Benchmark Control Problem
- A Benchmark Characterization of the EEMBC Benchmark Suite
- The Performance Validation of Linear Programming Algorithm Based on Integrated Benchmark
- An empirical assessment of Bellon's clone benchmark