A Primer in BERTology: What We Know About How BERT Works
Explore this paper's citation graph
Summary
This paper is the first survey of over 150 studies of the popular BERT model, reviewing the current state of knowledge about how BERT works, what kind of information it learns and how it is represented, common modifications to its training objectives and architecture, the overparameterization issue, and approaches to compression.
- Type
- preprint
- Published
- 2020-02-27
- Cited by
- 1,973
- References
- 208
- Access
- Open access
- OpenAlex
- https://openalex.org/W3006881356
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:211532403
Keywords
Transformer, Architecture, Computer science, State (computer science), Artificial intelligence
References
- Constructions at Work: The Nature of Generalization in Language
- Distilling the Knowledge in a Neural Network
- The Berkeley FrameNet Project
- Products of Random Latent Variable Grammars
- Distributed Representations of Words and Phrases and their Compositionality
- GloVe: Global Vectors for Word Representation
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- All-but-the-Top: Simple and Effective Postprocessing for Word Representations
- Questionable Answers in Question Answering Research: Reproducibility and Variability of Published Results
- Sentence Encoders on STILTs: Supplementary Training on Intermediate Labeled-data Tasks
- Analysis Methods in Neural Language Processing: A Survey
- Assessing BERT's Syntactic Abilities
- Learning and Evaluating General Linguistic Intelligence
- An Analysis of Encoder Representations in Transformer-Based Machine Translation
- Parameter-Efficient Transfer Learning for NLP
- Cross-lingual Language Model Pretraining
- Cloze-driven Pretraining of Self-attention Networks
- To Tune or Not to Tune? Adapting Pretrained Representations to Diverse Tasks
- Linguistic Knowledge and Transferability of Contextual Representations
- On Measuring Social Biases in Sentence Encoders
Cited by
- Survey on evaluation methods for dialogue systems
- Fixed Encoder Self-Attention Patterns in Transformer-Based Machine Translation
- Compressing Large-Scale Transformer-Based Models: A Case Study on BERT
- What the [MASK]? Making Sense of Language-Specific BERT Models
- Sentence Analogies: Exploring Linguistic Relationships and Regularities in Sentence Embeddings
- Pre-trained models for natural language processing: A survey
- Understanding Cross-Lingual Syntactic Transfer in Multilingual Recurrent Neural Networks
- Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence
- Telling BERT’s Full Story: from Local Attention to Global Aggregation
- DynaBERT: Dynamic BERT with Adaptive Width and Depth
- A Systematic Analysis of Morphological Content in BERT Models for Multiple Languages
- What’s so special about BERT’s layers? A closer look at the NLP pipeline in monolingual and multilingual models
- Cross-lingual Contextualized Topic Models with Zero-shot Learning
- Attention Module is Not Only a Weight: Analyzing Transformers with Vector Norms
- Syntactic Structure from Deep Learning
- Asking without Telling: Exploring Latent Ontologies in Contextual Representations
- Beneath the Tip of the Iceberg: Current Challenges and New Directions in Sentiment Analysis Research
- Enriched Pre-trained Transformers for Joint Slot Filling and Intent Detection
- How Do Decisions Emerge across Layers in Neural Models? Interpretation with Differentiable Masking
- What Happens To BERT Embeddings During Fine-tuning?