AI Research Atlas

Fable 5 cyber-safeguard details and jailbreak severity framework

Anthropic · 2 July 2026

Anthropic published what Fable 5's cyber classifiers do and do not block, plus a draft severity scale for AI jailbreaks built with Glasswing partners.

After redeployment Anthropic listed the harm types its cyber classifiers target and proposed a framework for grading jailbreak severity, so labs and governments can describe a bypass in shared terms. Developed with Amazon, Microsoft, Google and other Glasswing partners; presented as an early draft for discussion.

Date
Thursday, 2 July 2026
Lab
Anthropic
Kind
paper
Access
paper only

Only the post's opening was read in full; the framework's tiers and the classifier harm list were not transcribed here.

Sources

  1. www.anthropic.com/news/fable-safeguards-jailbreak-framework
  2. www.anthropic.com/news/redeploying-fable-5

This record was checked against its sources on 6 October 2026. How we check

Related