Let the Models Respond: Interpreting Language Model Detoxification Through the Lens of Prompt Dependence

Explore this paper's citation graph

Summary

This work applies popular detoxification approaches to several language models and quantifies their impact on the resulting models' prompt dependence using feature attribution methods, evaluating the effectiveness of counter-narrative fine-tuning and comparing it with reinforcement learning-driven detoxification.

Type
preprint
Published
2023-09-01
Cited by
0
References
30
Access
Open access

Keywords

Detoxification (alternative medicine), Computer science, Language model, Narrative, Attribution

References

Cited by

No citing papers recorded for this paper.

Related papers