Supplementary Material: Implementation and Experiments for GAU-based Model
Explore this paper's citation graph
Summary
A novel GAU-based model is proposed, which achieves a dev average score of 75.02, 1% higher than RoFormerV1 and being 45% faster, which is also competitive with Ro formerV2.
- Type
- preprint
- Published
- 2022-05-12
- Cited by
- 0
- References
- 24
- Access
- Open access
- OpenAlex
- https://openalex.org/W4280556182
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:248721714
Keywords
Computer science, Benchmark (surveying), Transformer, Footprint, Memory footprint
References
- Gaussian Error Linear Units (GELUs)
- Generating Long Sequences with Sparse Transformers
- Pre-Training with Whole Word Masking for Chinese BERT
- Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- GLU Variants Improve Transformer
- CLUE: A Chinese Language Understanding Evaluation Benchmark
- Longformer: The Long-Document Transformer
- Rethinking Attention with Performers
- Do Transformer Modifications Transfer Across Implementations and Applications?
- Finetuning Pretrained Transformers into RNNs
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- DeepNet: Scaling Transformers to 1,000 Layers
- Big Bird: Transformers for Longer Sequences
- GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
- Attention is All you Need
- Transformer Quality in Linear Time
- Primer: Searching for Efficient Transformers for Language Modeling
- cosFormer: Rethinking Softmax in Attention
- Language Models are Unsupervised Multitask Learners
Cited by
No citing papers recorded for this paper.
Related papers
- Theoretical Analysis of the Benchmark for Choosing Manipulative Instruments of Monetary Policies
- Exploring disk performance benchmarks
- Solutions to the Third Benchmark Control Problem
- A Benchmark Characterization of the EEMBC Benchmark Suite
- The Performance Validation of Linear Programming Algorithm Based on Integrated Benchmark
- An empirical assessment of Bellon's clone benchmark
- An Ensemble Deep Neural Network for Footprint Image Retrieval Based on Transfer Learning