AI Research Atlas

Constitutional Classifiers

Anthropic · 3 February 2025

Classifiers trained on synthetic data from a written constitution cut jailbreak success on Claude from 86% to 4.4% with minimal over-refusal.

Input and output classifiers trained on constitution-derived synthetic data screen jailbreaks. Over-refusals rise only 0.38% (not statistically significant); compute overhead is 23.7%. A public red-team demo (3-10 Feb 2025) drew 339 jailbreakers; four passed all eight levels, one with a universal jailbreak.

Date
Monday, 3 February 2025
Lab
Anthropic
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Jailbreak success rate4.4% (vs 86% unguarded)
Anthropic synthetic jailbreak test set
company
Compute overhead+23.7%company
Public demo participants339
about 300,000 chat interactions, 3,700 hours
company

Constitutional Classifiers were deployed with Claude Opus 4 under ASL-3 (2025-05-22). Numbers are from Anthropic's own synthetic jailbreak set; the live demo found a universal jailbreak.

Sources

  1. www.anthropic.com/news/constitutional-classifiers

This record was checked against its sources on 6 October 2026. How we check