Next-generation Constitutional Classifiers
Second-generation Constitutional Classifiers cut the extra compute for jailbreak defense from 23.7% to about 1% and lower refusals on harmless queries.
First-generation classifiers cut jailbreak success from 86% to 4.4% but cost 23.7% more compute and 0.38% more refusals. The new design keeps robustness against universal jailbreaks while being cheap enough to run on every request. Basis for the classifiers that later guard Fable and Opus models.
- Date
- Friday, 9 January 2026
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Extra compute for classifier defense | about 1% vs 23.7% for the first generation | company |
Anthropic's own robustness claims; the 2026 jailbreak fallout around Fable 5 shows classifier gaps remain.
Sources
This record was checked against its sources on 6 October 2026. How we check