Let the Models Respond: Interpreting Language Model Detoxification Through the Lens of Prompt Dependence
Explore this paper's citation graph
Summary
This work applies popular detoxification approaches to several language models and quantifies their impact on the resulting models' prompt dependence using feature attribution methods, evaluating the effectiveness of counter-narrative fine-tuning and comparing it with reinforcement learning-driven detoxification.
- Type
- preprint
- Published
- 2023-09-01
- Cited by
- 0
- References
- 30
- Access
- Open access
- OpenAlex
- https://openalex.org/W4386499048
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:261530972
Keywords
Detoxification (alternative medicine), Computer science, Language model, Narrative, Attribution
References
- CONAN - COunter NArratives through Nichesourcing: a Multilingual Dataset of Responses to Fight Online Hate Speech
- RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
- Analyzing the Source and Target Contributions to Predictions in Neural Machine Translation
- Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Entropy-based Attention Regularization Frees Unintended Bias Mitigation from Lists
- Using Pre-Trained Language Models for Producing Counter Narratives Against Hate Speech: a Comparative Study
- Towards Opening the Black Box of Neural Machine Translation: Source and Target Interpretations of the Transformer
- Taxonomy of Risks posed by Language Models
- No Language Left Behind: Scaling Human-Centered Machine Translation
- Improving alignment of dialogue agents via targeted human judgements
- Human-Machine Collaboration Approaches to Build a Dialogue Dataset for Hate Speech Countering
- Constitutional AI: Harmlessness from AI Feedback
- Pretraining Language Models with Human Preferences
- Explaining How Transformers Use Context to Build Predictions
- Jailbroken: How Does LLM Safety Training Fail?
- Inseq: An Interpretability Toolkit for Sequence Generation Models
- XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
- LIMA: Less Is More for Alignment
- Training language models to follow instructions with human feedback
Cited by
No citing papers recorded for this paper.
Related papers
- Domiciliary detoxification: a cost effective alternative to inpatient treatment.
- Alternative strategies of opiate detoxification: Evaluation of the so-called ultra rapid detoxification
- [Detoxification of patients with GHB dependence].
- Methadone Maintenance to Abstinence: How Many Make It?
- Similar efficacy of abrupt and gradual opiate detoxification.
- Detoxification from alcohol: a comparison of home detoxification and hospital-based day patient care.
- A 24-h inpatient detoxification treatment for heroin addicts: a preliminary investigation.
- Detoxification of rehabilitated methadone patients: frequency and predictors of long-term success.