Locating and Editing Factual Associations in GPT
Explore this paper's citation graph
Summary
An important role for mid-layer feed-forward modules in storing factual associations in autoregressive transformer language models is confirmed and direct manipulation of computational mechanisms may be a feasible approach for model editing.
- Type
- preprint
- Published
- 2022-02-10
- Cited by
- 3,128
- References
- 59
- Access
- Open access
- OpenAlex
- https://openalex.org/W4281657280
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:255825985
Keywords
Computer science, Counterfactual thinking, Transformer, Generalization, Task (project management)
References
- A simple neural network generating an interactive memory
- Direct and Indirect Effects
- Correlation Matrix Memories
- Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
- Probing for semantic evidence of composition by means of simple classification tasks
- Axiomatic Attribution for Deep Networks
- What do Neural Machine Translation Models Learn about Morphology?
- Zero-Shot Relation Extraction via Reading Comprehension
- Analysis Methods in Neural Language Processing: A Survey
- Generating Informative and Diverse Conversational Responses via Adversarial Information Maximization
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
- Language Models as Knowledge Bases?
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
- How Can We Know What Language Models Know?
- How Much Knowledge Can You Pack into the Parameters of a Language Model?
- Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias
- How Context Affects Language Models' Factual Predictions
- CausaLM: Causal Model Explanation Through Counterfactual Language Models
Cited by
- The alignment problem from a deep learning perspective
- Extremely Simple Activation Shaping for Out-of-Distribution Detection
- Learning by Distilling Context
- Mass-Editing Memory in a Transformer
- Language Generation Models Can Cause Harm: So What Can We Do About It? An Actionable Survey
- Revision Transformers: Getting RiT of No-Nos
- Causal Analysis of Syntactic Agreement Neurons in Multilingual Language Models
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
- On the Domain Adaptation and Generalization of Pretrained Language Models: A Survey
- DisentQA: Disentangling Parametric and Contextual Knowledge with Counterfactual Question Answering
- Diagnostics for Deep Neural Networks with Automated Copy/Paste Attacks
- Convexifying Transformers: Improving optimization and understanding of transformer networks
- Language Models as Agent Models
- When Neural Model Meets NL2Code: A Survey
- DSI++: Updating Transformer Memory with New Documents
- DialGuide: Aligning Dialogue Model Behavior with Developer Guidelines
- A Survey on Knowledge-Enhanced Pre-trained Language Models
- Tracr: Compiled Transformers as a Laboratory for Interpretability
- Augmented Behavioral Annotation Tools, with Application to Multimodal Datasets and Models: A Systematic Review
- Analyzing Feed-Forward Blocks in Transformers through the Lens of Attention Maps
Related papers
- Counterfactual Instances Explain Little
- Ai-based counterfactual reasoning for tourism research
- Meta-supervision for Attention Using Counterfactual Estimation
- Making inferences between counterfactual and real worlds: developmental evidence
- Experimental Study of Characteristics of Counterfactual Thinking in Oldmen Individuals
- Use Case of Counterfactual Examples: Data Augmentation
- 3 to 5 Year-old Children’s Acting on Antecedent Counterfactual Reasoning
- Empowering Language Understanding with Counterfactual Reasoning