AI Research Atlas

Responsible Scaling Policy update (v2)

Anthropic · 15 October 2024

Anthropic's October 2024 RSP update adds capability thresholds that trigger stronger safeguards and safety-case-style evaluation of those safeguards.

More flexible than v1: defined thresholds for when to upgrade safeguards, refined processes to evaluate capabilities and safeguard adequacy (inspired by safety cases), and new internal governance and external input, while keeping the commitment not to train or deploy without adequate safeguards.

Date
Tuesday, 15 October 2024
Lab
Anthropic
Kind
paper
Access
paper only

Superseded by v3 on 2026-02-24. Its ASL-3 standard was first applied to Claude Opus 4 in May 2025.

Sources

  1. www.anthropic.com/news/announcing-our-updated-responsible-scaling-policy

This record was checked against its sources on 6 October 2026. How we check

Related