Jailbroken: How Does LLM Safety Training Fail?

Explore this paper's citation graph

Summary

The analysis emphasizes the need for safety-capability parity -- that safety mechanisms should be as sophisticated as the underlying model -- and argues against the idea that scaling alone can resolve these safety failure modes.

Type
preprint
Published
2023-07-05
Cited by
2,121
References
67
Access
Open access

Keywords

Adversarial system, Generalization, Computer security, Computer science, Training (meteorology)

References

Cited by

Related papers