Responsible Scaling Policy v1
Anthropic publishes its Responsible Scaling Policy, which sets capability-triggered AI Safety Levels (ASL) with required safeguards before further scaling or deployment.
A capability-threshold safety policy modeled on biosafety levels. ASL-2 covers current models; ASL-3 requires unusually strong security and a commitment not to deploy models that show meaningful catastrophic-misuse risk under adversarial testing. Training can pause if safeguards lag. Board approval and Long Term Benefit Trust consultation needed to change it.
- Date
- Tuesday, 19 September 2023
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Policy document, not a model. Its ASL-3 tier was first invoked for Claude Opus 4 on 2025-05-22 (see anthropic-asl3-activation).
Sources
This record was checked against its sources on 6 October 2026. How we check