Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Explore this paper's citation graph
Summary
This systematic study compares pre-training objectives, architectures, unlabeled datasets, transfer approaches, and other factors on dozens of language understanding tasks and achieves state-of-the-art results on many benchmarks covering summarization, question answering, text classification, and more.
- Type
- preprint
- Published
- 2019-10-23
- Cited by
- 27,088
- References
- 134
- Access
- Open access
- OpenAlex
- https://openalex.org/W2981852735
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:204838007
Keywords
Automatic summarization, Computer science, Transfer of learning, Natural language processing, Artificial intelligence
References
- Automatically Constructing a Corpus of Sentential Paraphrases
- Teaching Machines to Read and Comprehend
- Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books
- One weird trick for parallelizing convolutional neural networks
- Efficient Estimation of Word Representations in Vector Space
- Resolving Complex Cases of Definite Pronouns: The Winograd Schema Challenge
- Generating Sequences With Recurrent Neural Networks
- Neural Machine Translation of Rare Words with Subword Units
- Distilling the Knowledge in a Neural Network
- A Learning Algorithm for Continually Running Fully Recurrent Neural Networks
- A bitter lesson.
- Dropout: a simple way to prevent neural networks from overfitting
- Bleu: a Method for Automatic Evaluation of Machine Translation
- ImageNet: A large-scale hierarchical image database
- Federated Optimization: Distributed Optimization Beyond the Datacenter
- ImageNet Large Scale Visual Recognition Challenge
- A Convolutional Neural Network for Modelling Sentences
- Choice of Plausible Alternatives: An Evaluation of Commonsense Causal Reasoning
- Distributed Representations of Words and Phrases and their Compositionality
- ROUGE: A Package for Automatic Evaluation of Summaries
Cited by
- ChatGPT makes medicine easy to swallow: an exploratory case study on simplified radiology reports
- Learning entity-oriented representation for biomedical relation extraction
- Can Language Models Trained on Written Monologue Learn to Predict Spoken Dialogue?
- Deep Learning for Genomics: A Concise Overview
- Contextual Word Representations: A Contextual Introduction
- Are All Layers Created Equal?
- Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
- Taming Momentum in a Distributed Asynchronous Environment
- c-TextGen: Conditional Text Generation for Harmonious Human-Machine Interaction
- Direct-fit to nature: an evolutionary perspective on biological (and artificial) neural networks
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- TinyBERT: Distilling BERT for Natural Language Understanding
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- Portuguese Named Entity Recognition using BERT-CRF
- Transformers: State-of-the-Art Natural Language Processing
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
- Discourse-Aware Neural Extractive Model for Text Summarization
- Multi-Stage Document Ranking with BERT
- INSET: Sentence Infilling with Inter-sentential Generative Pre-training
Related papers
- Language Models are Few-Shot Learners
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding