AI Research Atlas

How we contain Claude across products

Anthropic · 25 May 2026

Anthropic explains capping an agent's 'blast radius' through environment controls, after telemetry showed users approved about 93% of permission prompts.

Argues failure likelihood has fallen but potential damage only grows, so containment (sandboxes, limited access) must complement human-in-the-loop prompts that users rubber-stamp. Cites Mythos Preview as a model whose blast radius was judged too high to ship in April 2026.

Date
Monday, 25 May 2026
Lab
Anthropic
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Permission prompts approved by Claude Code usersabout 93%
Anthropic telemetry
company

Published two months before the cybersecurity-evaluation incidents Anthropic disclosed on 2026-07-30.

Sources

  1. www.anthropic.com/engineering/how-we-contain-claude

This record was checked against its sources on 6 October 2026. How we check

Related