BERT Rediscovers the Classical NLP Pipeline
Explore this paper's citation graph
Summary
This work finds that the model represents the steps of the traditional NLP pipeline in an interpretable and localizable way, and that the regions responsible for each step appear in the expected sequence: POS tagging, parsing, NER, semantic roles, then coreference.
- Type
- article
- Published
- 2019-05-15
- Cited by
- 2,032
- References
- 22
- Access
- Open access
- OpenAlex
- https://openalex.org/W2946417913
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:155092004
Keywords
Coreference, Pipeline (software), Computer science, Natural language processing, Parsing
References
- SemEval-2010 Task 8: Multi-Way Classification of Semantic Relations Between Pairs of Nominals
- The Stanford CoreNLP Natural Language Processing Toolkit
- Distributed Representations of Words and Phrases and their Compositionality
- Semantic Proto-Roles
- A Gold Standard Dependency Corpus for English
- Does String-Based Neural MT Learn Source Syntax?
- Semantic Proto-Role Labeling
- Deep Contextualized Word Representations
- Deep RNNs Encode Soft Hierarchical Syntax
- Dissecting Contextual Word Embeddings: Architecture and Representation
- Targeted Syntactic Evaluation of Language Models
- On internal language representations in deep learning: an analysis of machine translation and speech recognition
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Multi-Task Deep Neural Networks for Natural Language Understanding
- What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
- MizAR 60 for Mizar 50
- Attention is All you Need
- What do you learn from context? Probing for sentence structure in contextualized word representations
- Language Models are Unsupervised Multitask Learners
- Scholarship, Research, and Creative Work at Bryn Mawr College Scholarship, Research, and Creative Work at Bryn Mawr College
Cited by
- Learning entity-oriented representation for biomedical relation extraction
- Toward Fast and Accurate Neural Chinese Word Segmentation with Multi-Criteria Learning
- How Multilingual is Multilingual BERT?
- Visualizing and Measuring the Geometry of BERT
- Analyzing the Structure of Attention in a Transformer Language Model
- Theoretical Limitations of Self-Attention in Neural Sequence Models
- What Does BERT Look at? An Analysis of BERT’s Attention
- X-BERT: eXtreme Multi-label Text Classification with BERT
- Leveraging Pre-trained Checkpoints for Sequence Generation Tasks
- What BERT Is Not: Lessons from a New Suite of Psycholinguistic Diagnostics for Language Models
- Visualizing and Understanding the Effectiveness of BERT
- The compositionality of neural networks: integrating symbolism and connectionism
- Learning Latent Parameters without Human Response Patterns: Item Response Theory with Artificial Crowds
- Unsupervised Labeled Parsing with Deep Inside-Outside Recursive Autoencoders
- Evaluating BERT for natural language inference: A case study on the CommitmentBank
- 75 Languages, 1 Model: Parsing Universal Dependencies Universally
- Adaptively Sparse Transformers
- The Bottom-up Evolution of Representations in the Transformer: A Study with Machine Translation and Language Modeling Objectives
- Does BERT agree? Evaluating knowledge of structure dependence through agreement relations
- Effective Use of Transformer Networks for Entity Tracking
Related papers
- End-to-end coreference resolution for clinical narratives
- DialogRE^C+: An Extension of DialogRE to Investigate How Much Coreference Helps Relation Extraction in Dialogs
- Constrained Multi-Task Learning for Event Coreference Resolution
- Impact of Coreference Resolution on Slot Filling
- A Unified Event Coreference Resolution by Integrating Multiple Resolvers
- Research on the Feature Selection of Chinese Coreference Resolution
- Towards Harnessing Memory Networks for Coreference Resolution
- What is coreference, and what should coreference annotation be?