AI Research Atlas

Many-shot jailbreaking

Anthropic · 2 April 2024

Many-shot jailbreaking fills a long context window with fake harmful Q&A turns to override safety training, and works better as context grows.

Shows long context windows, then up to 1M tokens, create a new attack surface. Effective on Anthropic's and other labs' models; Anthropic briefed other developers first and added mitigations before publishing.

Date
Tuesday, 2 April 2024
Lab
Anthropic
Kind
paper
Access
paper only

Attack-success curves are in the paper and were not transcribed here. Inferred to be part of the lineage leading to the 2025 Constitutional Classifiers.

Sources

  1. www.anthropic.com/research/many-shot-jailbreaking

This record was checked against its sources on 6 October 2026. How we check

Related