Many-shot jailbreaking
Many-shot jailbreaking fills a long context window with fake harmful Q&A turns to override safety training, and works better as context grows.
Shows long context windows, then up to 1M tokens, create a new attack surface. Effective on Anthropic's and other labs' models; Anthropic briefed other developers first and added mitigations before publishing.
- Date
- Tuesday, 2 April 2024
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Attack-success curves are in the paper and were not transcribed here. Inferred to be part of the lineage leading to the 2025 Constitutional Classifiers.
Sources
This record was checked against its sources on 6 October 2026. How we check