Measuring Massive Multitask Language Understanding

Explore this paper's citation graph

Summary

While most recent models have near random-chance accuracy, the very largest GPT-3 model improves over random chance by almost 20 percentage points on average, however, on every one of the 57 tasks, the best models still need substantial improvements before they can reach expert-level accuracy.

Type
preprint
Published
2020-09-07
Cited by
9,327
References
35
Access
Open access

Keywords

Computer science, Test (biology), Artificial intelligence, Machine learning, Measure (data warehouse)

References

Cited by

Related papers