GLU Variants Improve Transformer
Explore this paper's citation graph
Summary
Gated Linear Units (GLU) consist of the component-wise product of two linear projections, one of which is first passed through a sigmoid function, and it is found that some of them yield quality improvements over the typically-used ReLU or GELU activations.
- Type
- preprint
- Published
- 2020-02-12
- Cited by
- 2,186
- References
- 12
- Access
- Open access
- OpenAlex
- https://openalex.org/W3006439205
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:211096588
Keywords
Sigmoid function, Transformer, Nonlinear system, Mathematics, Sequence (biology)
References
- Three new graphical models for statistical language modelling
- Deep Sparse Rectifier Neural Networks
- Bridging Nonlinearities and Stochastic Regularizers with Gaussian Error Linear Units
- Searching for Activation Functions
- GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
- SQuAD: 100,000+ Questions for Machine Comprehension of Text
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
- Attention is All you Need
- Adafactor: Adaptive Learning Rates with Sublinear Memory Cost
- GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
- Language Modeling with Gated Convolutional Networks
Cited by
- How Much Knowledge Can You Pack into the Parameters of a Language Model?
- mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer
- Do Transformer Modifications Transfer Across Implementations and Applications?
- The Power of Scale for Parameter-Efficient Prompt Tuning
- Memory-efficient Transformers via Top-k Attention
- A Survey of Transformers
- Lightweight Causal Transformer with Local Self-Attention for Real-Time Speech Enhancement
- Online Compressive Transformer for End-to-End Speech Recognition
- Memory transformer with hierarchical attention for long document processing
- IAGC: Interactive Attention Graph Convolution Network for Semantic Segmentation of Point Clouds in Building Indoor Environment
- Geometry-Aware Supertagging with Heterogeneous Dynamic Convolutions
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
- Simple Baselines for Image Restoration
- What Language Model Architecture and Pretraining Objective Work Best for Zero-Shot Generalization?
- Boosting Adversarial Transferability of MLP-Mixer
- Error Correction Code Transformer
- Supplementary Material: Implementation and Experiments for GAU-based Model
- Life after BERT: What do Other Muppets Understand about Language?
- BanglaNLG: Benchmarks and Resources for Evaluating Low-Resource Natural Language Generation in Bangla
- Sparse Mixers: Combining MoE and Mixing to build a more efficient BERT
Related papers
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding