Large Language Models are Zero-Shot Reasoners
Explore this paper's citation graph
Summary
Experimental results demonstrate that the Zero-shot-CoT, using the same single prompt template, significantly outperforms zero-shot LLM performances on diverse benchmark reasoning tasks including arithmetics, symbolic reasoning, and other logical reasoning tasks, without any hand-crafted few-shot examples.
- Type
- preprint
- Published
- 2022-05-24
- Cited by
- 7,895
- References
- 61
- Access
- Open access
- OpenAlex
- https://openalex.org/W4281557260
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:249017743
Keywords
Shot (pellet), Task (project management), Benchmark (surveying), Computer science, Zero (linguistics)
References
- A measure of intelligence
- The structure of human intelligence: It is verbal, perceptual, and image rotation (VPR), not fluid and crystallized
- Learning to Solve Arithmetic Word Problems with Verb Categorization
- Solving General Arithmetic Word Problems
- Parsing Algebraic Word Problems into Equations
- The Cattell-Horn-Carroll Theory of Cognitive Abilities: Past, Present, and Future.
- MAWPS: A Math Word Problem Repository
- Explain Yourself! Leveraging Language Models for Commonsense Reasoning
- Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Unsupervised Commonsense Question Answering with Self-Talk
- 5分で分かる!? 有名論文ナナメ読み:Jacob Devlin et al. : BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding
- It’s Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners
- What Makes Good In-Context Examples for GPT-3?
- Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm
- Are NLP Models really able to Solve Simple Math Word Problems?
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity
- Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies
Cited by
- Statistically Profiling Biases in Natural Language Reasoning Datasets and Models
- Artificial intelligence in systematic literature reviews: A case for cautious optimism.
- Generating Natural Language Proofs with Verifier-Guided Search
- On the Advance of Making Language Models Better Reasoners
- A Dataset and Benchmark for Automatically Answering and Generating Machine Learning Final Exams
- Emergent Abilities of Large Language Models
- LIFT: Language-Interfaced Fine-Tuning for Non-Language Machine Learning Tasks
- MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge
- Repository-Level Prompt Generation for Large Language Models of Code
- Automatic Generation of Programming Exercises and Code Explanations Using Large Language Models
- Rationale-Augmented Ensembles in Language Models
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- Language models show human-like content effects on reasoning
- Targeted Honeyword Generation with Language Models
- World robot challenge 2020 – partner robot: a data-driven approach for room tidying with mobile manipulator
- FOLIO: Natural Language Reasoning with First-Order Logic
- Accelerating Antibody Design with Active Learning
- Do Androids Laugh at Electric Sheep? Humor “Understanding” Benchmarks from The New Yorker Caption Contest
- Dynamic Generation of Interpretable Inference Rules in a Neuro-Symbolic Expert System
- AutoPET Challenge 2022: Step-by-Step Lesion Segmentation in Whole-body FDG-PET/CT
Related papers
- High-speed photographic study on shot put
- Study on viewer’s preference of sensibility vocabulary depending on composition of portrait shot in image -Mainly on the basis of size and placement of shot-
- Study on Shot and be Shot Skills of Our Men Soccer Player from the World Cup
- Theoretical Analysis of the Benchmark for Choosing Manipulative Instruments of Monetary Policies
- A shot-by-shot breakdown of Remembrance
- A shot-by-shot breakdown of In Chambers
- Exploring disk performance benchmarks
- Influence of Shot Glissading Skill on Shot Put Distance of Male College students