Scaling Instruction-Finetuned Language Models
Explore this paper's citation graph
Summary
It is found that instruction finetuning with the above aspects dramatically improves performance on a variety of model classes (PaLM, T5, U-PaLM), prompting setups, and evaluation benchmarks (MMLU, BBH, TyDiQA, MGSM, open-ended generation).
- Type
- preprint
- Published
- 2022-10-20
- Cited by
- 4,394
- References
- 106
- Access
- Open access
- OpenAlex
- https://openalex.org/W4307079201
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:253018554
Keywords
Computer science, Margin (machine learning), Usability, Variety (cybernetics), Scaling
References
- Anticipatory Ethics for Emerging Technologies
- The LAMBADA dataset: Word prediction requiring a broad discourse context
- Gender Bias in Coreference Resolution
- e-SNLI: Natural Language Inference with Natural Language Explanations
- Model Cards for Model Reporting
- Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification
- Does it Make Sense? And Why? A Pilot Study for Sense Making and Explanation
- Explain Yourself! Leveraging Language Models for Commonsense Reasoning
- Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Taskmaster-1: Toward a Realistic and Diverse Dialog Dataset
- Make Up Your Mind! Adversarial Generation of Inconsistent Natural Language Explanations
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- QASC: A Dataset for Question Answering via Sentence Composition
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
- SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
- Welcome, Singular "They"
- Scaling Laws for Neural Language Models
- TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages
- UnifiedQA: Crossing Format Boundaries With a Single QA System
Cited by
- Emergent Abilities of Large Language Models
- MVP: Multi-task Supervised Pre-training for Natural Language Generation
- ThinkSum: Probabilistic reasoning over sets using large language models
- Recitation-Augmented Language Models
- Efficiently Enhancing Zero-Shot Performance of Instruction Following Model via Retrieval of Soft Prompt
- Guess the Instruction! Flipped Learning Makes Language Models Stronger Zero-Shot Learners
- Transcending Scaling Laws with 0.1% Extra Compute
- Help me write a Poem - Instruction Tuning as a Vehicle for Collaborative Poetry Writing
- CONDAQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation
- Two-stage LLM Fine-tuning with Less Specialization and More Generalization
- A Universal Discriminator for Zero-Shot Generalization
- HyperTuning: Toward Adapting Large Language Models without Back-propagation
- GPT-Neo for commonsense reasoning-a theoretical and practical lens
- ColD Fusion: Collaborative Descent for Distributed Multitask Finetuning
- CLAM: Selective Clarification for Ambiguous Questions with Large Language Models
- On Second Thought, Let’s Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning
- Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models
- Teaching Small Language Models to Reason
- Improving Cross-task Generalization of Unified Table-to-text Models with Compositional Task Configurations
- Can Retriever-Augmented Language Models Reason? The Blame Game Between the Retriever and the Language Model
Related papers
- Enhancing the Human-Machine Interface: An Analysis of Unmanned Aerial Systems (UAS) Using the System Usability Scale (SUS)
- Development of Usability Assessment Indicator of Adjustable Electric Beds for Home Care
- Making energy savings easier: usability metrics for thermostats
- Diagnosis and identification of key issues of usability for reducing medication errors
- Usability assessment of a mobile app for patients with peripherally inserted central catheters
- Usability Evaluation of an Admission, Discharge, and Transfer Information System: A Heuristic Evaluation
- Usability evaluation of ambient assisted living systems using a multi-method approach
- A longitudinal study of usability in health care - does time heal?