AI Research Atlas

Next-generation Constitutional Classifiers

Anthropic · 9 January 2026

Second-generation Constitutional Classifiers cut the extra compute for jailbreak defense from 23.7% to about 1% and lower refusals on harmless queries.

First-generation classifiers cut jailbreak success from 86% to 4.4% but cost 23.7% more compute and 0.38% more refusals. The new design keeps robustness against universal jailbreaks while being cheap enough to run on every request. Basis for the classifiers that later guard Fable and Opus models.

Date
Friday, 9 January 2026
Lab
Anthropic
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Extra compute for classifier defenseabout 1%
vs 23.7% for the first generation
company

Anthropic's own robustness claims; the 2026 jailbreak fallout around Fable 5 shows classifier gaps remain.

Sources

  1. www.anthropic.com/research/next-generation-constitutional-classifiers

This record was checked against its sources on 6 October 2026. How we check

Related