Jailbroken: How Does LLM Safety Training Fail?
Explore this paper's citation graph
Summary
The analysis emphasizes the need for safety-capability parity -- that safety mechanisms should be as sophisticated as the underlying model -- and argues against the idea that scaling alone can resolve these safety failure modes.
- Type
- preprint
- Published
- 2023-07-05
- Cited by
- 2,121
- References
- 67
- Access
- Open access
- OpenAlex
- https://openalex.org/W4383473937
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:259342528
Keywords
Adversarial system, Generalization, Computer security, Computer science, Training (meteorology)
References
- SAFETY THROUGH DESIGN
- Adversarial Attacks and Defences: A Survey
- Universal Adversarial Triggers for Attacking and Analyzing NLP
- Fine-Tuning Language Models from Human Preferences
- RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
- Recipes for Safety in Open-domain Chatbots
- Extracting Training Data from Large Language Models
- All the News That’s Fit to Fabricate: AI-Generated Text as a Tool of Media Misinformation
- Challenges in Detoxifying Language Models
- Style.
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Emergent Abilities of Large Language Models
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- On the Impossible Safety of Large AI Models
- On Second Thought, Let’s Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning
- Constitutional AI: Harmlessness from AI Feedback
- Generative Language Models and Automated Influence Operations: Emerging Threats and Potential Mitigations
- Pretraining Language Models with Human Preferences
- More than you've asked for: A Comprehensive Analysis of Novel Prompt Injection Threats to Application-Integrated Large Language Models
- Automatically Auditing Large Language Models via Discrete Optimization
Cited by
- Explore, Establish, Exploit: Red Teaming Language Models from Scratch
- Visual Adversarial Examples Jailbreak Aligned Large Language Models
- Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language Models
- LLM Censorship: A Machine Learning Challenge or a Computer Security Problem?
- Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models
- AgentBench: Evaluating LLMs as Agents
- GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
- LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
- Robustness Over Time: Understanding Adversarial Examples’ Effectiveness on Longitudinal Versions of Large Language Models
- XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
- Using Large Language Models for Cybersecurity Capture-The-Flag Challenges and Certification Questions
- Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions
- Sparks of Large Audio Models: A Survey and Outlook
- Use of LLMs for Illicit Purposes: Threats, Prevention Measures, and Vulnerabilities
- Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
- Image Hijacks: Adversarial Images can Control Generative Models at Runtime
- Let the Models Respond: Interpreting Language Model Detoxification Through the Lens of Prompt Dependence
- Certifying LLM Safety against Adversarial Prompting
- Open Sesame! Universal Black Box Jailbreaking of Large Language Models
- Down the Toxicity Rabbit Hole: Investigating PaLM 2 Guardrails