Constitutional Classifiers
Classifiers trained on synthetic data from a written constitution cut jailbreak success on Claude from 86% to 4.4% with minimal over-refusal.
Input and output classifiers trained on constitution-derived synthetic data screen jailbreaks. Over-refusals rise only 0.38% (not statistically significant); compute overhead is 23.7%. A public red-team demo (3-10 Feb 2025) drew 339 jailbreakers; four passed all eight levels, one with a universal jailbreak.
- Date
- Monday, 3 February 2025
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Jailbreak success rate | 4.4% (vs 86% unguarded) Anthropic synthetic jailbreak test set | company |
| Compute overhead | +23.7% | company |
| Public demo participants | 339 about 300,000 chat interactions, 3,700 hours | company |
Constitutional Classifiers were deployed with Claude Opus 4 under ASL-3 (2025-05-22). Numbers are from Anthropic's own synthetic jailbreak set; the live demo found a universal jailbreak.
Sources
This record was checked against its sources on 6 October 2026. How we check