GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Explore this paper's citation graph
Summary
A benchmark of nine diverse NLU tasks, an auxiliary dataset for probing models for understanding of specific linguistic phenomena, and an online platform for evaluating and comparing models, which favors models that can represent linguistic knowledge in a way that facilitates sample-efficient learning and effective knowledge-transfer across tasks.
- Type
- preprint
- Published
- 2018-04-20
- Cited by
- 8,891
- References
- 77
- Access
- Open access
- OpenAlex
- https://openalex.org/W2799054028
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:5034059
Keywords
Computer science, Natural language understanding, Benchmark (surveying), Task (project management), Suite
References
- Automatically Constructing a Corpus of Sentential Paraphrases
- Machine Learning Challenges. Evaluating Predictive Uncertainty, Visual Object Classification, and Recognising Tectual Entailment
- One billion word benchmark for measuring progress in statistical language modeling
- Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books
- A large annotated corpus for learning natural language inference
- Annotating Expressions of Opinions and Emotions in Language
- The empirical base of linguistics: Grammaticality judgments and linguistic methodology
- Comparing two K-category assignments by a K-category correlation coefficient
- Comparison of the predicted and observed secondary structure of T4 phage lysozyme.
- A Sentimental Education: Sentiment Analysis Using Subjectivity Summarization Based on Minimum Cuts
- Reasoning about Entailment with Neural Attention
- Distributed Representations of Sentences and Documents
- Mining and summarizing customer reviews
- Seeing Stars: Exploiting Class Relationships for Sentiment Categorization with Respect to Rating Scales
- Using the Framework
- GloVe: Global Vectors for Word Representation
- Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
- Learning Distributed Representations of Sentences from Unlabelled Data
- The Seventh PASCAL Recognizing Textual Entailment Challenge
- Bag of Tricks for Efficient Text Classification
Cited by
- Survey on evaluation methods for dialogue systems
- Neural Network Acceptability Judgments
- Does the brain represent words? An evaluation of brain decoding studies of language understanding
- The Natural Language Decathlon: Multitask Learning as Question Answering
- Trick Me If You Can: Adversarial Writing of Trivia Challenge Questions
- Recent Trends in Deep Learning Based Natural Language Processing
- A Survey of the Usages of Deep Learning for Natural Language Processing
- SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference
- Lexicosyntactic Inference in Neural Models
- Collecting Diverse Natural Language Inference Problems for Sentence Representation Evaluation
- Towards Text Generation with Adversarially Learned Neural Outlines
- XNLI: Evaluating Cross-lingual Sentence Representations
- Semi-Supervised Multi-Task Word Embeddings
- Language Modeling Teaches You More Syntax than Translation Does: Lessons Learned Through Auxiliary Task Analysis
- Sentence Encoders on STILTs: Supplementary Training on Intermediate Labeled-data Tasks
- Learning Distributed Representations of Symbolic Structure Using Binding and Unbinding Operations
- InferLite: Simple Universal Sentence Representations from Natural Language Inference Data
- Dialogue Natural Language Inference
- Combining Axiom Injection and Knowledge Base Completion for Efficient Natural Language Inference
- Non-entailed subsequences as a challenge for natural language inference
Related papers
- Perspector: Benchmarking Benchmark Suites
- The SPEC OMP2001 Benchmark on the Fujitsu PRIMEPOWER System
- A Benchmark programming assignment suite for quantitative analysis of student performance in early programming courses
- The Certimark benchmark: architecture and future perspectives
- A Benchmark-Suite of real-World constrained multi-objective optimization problems and some baseline results