Measuring Massive Multitask Language Understanding
Explore this paper's citation graph
Summary
While most recent models have near random-chance accuracy, the very largest GPT-3 model improves over random chance by almost 20 percentage points on average, however, on every one of the 57 tasks, the best models still need substantial improvements before they can reach expert-level accuracy.
- Type
- preprint
- Published
- 2020-09-07
- Cited by
- 9,327
- References
- 35
- Access
- Open access
- OpenAlex
- https://openalex.org/W3083410900
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:221516475
Keywords
Computer science, Test (biology), Artificial intelligence, Machine learning, Measure (data warehouse)
References
- Computing Machinery and Intelligence
- MCTest: A Challenge Dataset for the Open-Domain Machine Comprehension of Text
- The Arcade Learning Environment: An Evaluation Platform for General Agents
- The Arcade Learning Environment: An Evaluation Platform for General Agents (Extended Abstract)
- RACE: Large-scale ReAding Comprehension Dataset From Examinations
- On Calibration of Modern Neural Networks
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
- Deep Anomaly Detection with Outlier Exposure
- GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
- HellaSwag: Can a Machine Really Finish Your Sentence?
- Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift
- Natural Adversarial Examples
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Language Models as Knowledge Bases?
- Cosmos QA: Machine Reading Comprehension with Contextual Commonsense Reasoning
- Verified Uncertainty Calibration
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Cited by
- Measuring the Occupational Impact of AI: Tasks, Cognitive Abilities and AI Benchmarks
- Current Limitations of Language Models: What You Need is Retrieval
- Language Models are Open Knowledge Graphs
- Findings of the WMT 2023 Shared Task on Automatic Post-Editing
- What Makes Good In-Context Examples for GPT-3?
- Fairness for Unobserved Characteristics: Insights from Technological Impacts on Queer Communities
- Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm
- Limitations of Autoregressive Models and Their Alternatives
- Ethical-Advice Taker: Do Language Models Understand Natural Language Interventions?
- When does pretraining help?: assessing self-supervised learning for law and the CaseHOLD dataset of 53,000+ legal holdings
- Adapting Language Models for Zero-shot Learning by Meta-tuning on Dataset and Prompt Collections
- How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts
- Symbolic Knowledge Distillation: from General Language Models to Commonsense Models
- Few-Shot Self-Rationalization with Natural Language Prompts
- CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model Capabilities
- NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning Tasks
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- GPT-NeoX-20B: An Open-Source Autoregressive Language Model
Related papers
- An Improved Coronary Heart Disease Predictive System Using Random Forest
- Prediction and Characteristic Exploration of Military Specialized High School Trainee Selection Using Machine Learning
- Guided Random Forest in the RRF Package
- Enriched Random Forest for High Dimensional Genomic Data
- Comparing the Accuracy and Developed Models for Predicting the Confrontation Naming of the Elderly in South Korea using Weighted Random Forest, Random Forest, and Support Vector Regression
- A guided random forest based feature selection approach for activity recognition
- Diabetes Prediction using Random Forest Classifier with Different Wrapper Methods
- Review: William Shakespeare: Measure for Measure * Kate Chedgzoy: William Shakespeare: Measure for Measure
- An Improvised Random Forest Model for Breast Cancer Classification
- Implementation of LightGBM and Random Forest in Potential Customer Classification