Fable 5 cyber-safeguard details and jailbreak severity framework
Anthropic published what Fable 5's cyber classifiers do and do not block, plus a draft severity scale for AI jailbreaks built with Glasswing partners.
After redeployment Anthropic listed the harm types its cyber classifiers target and proposed a framework for grading jailbreak severity, so labs and governments can describe a bypass in shared terms. Developed with Amazon, Microsoft, Google and other Glasswing partners; presented as an early draft for discussion.
- Date
- Thursday, 2 July 2026
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Only the post's opening was read in full; the framework's tiers and the classifier harm list were not transcribed here.
Sources
- www.anthropic.com/news/fable-safeguards-jailbreak-framework
- www.anthropic.com/news/redeploying-fable-5
This record was checked against its sources on 6 October 2026. How we check