AI Research Atlas

Claude Code auto mode

Anthropic · 24 March 2026

Claude Code auto mode lets classifiers approve routine actions, with a 0.4% false-positive rate on real traffic but a 17% miss rate on overeager actions.

A middle path between per-action approval and no guardrails. A server-side probe screens tool outputs for prompt injection and a reasoning-blind transcript classifier (Sonnet 4.6) judges each action in two stages. Research preview on Team 2026-03-24, generally available 2026-07-10, default for new Pro, Max and Team sessions from 2026-08-14.

Date
Tuesday, 24 March 2026
Lab
Anthropic
Kind
feature
Access
closed API

Figures

MeasureValueMeasured by
False-positive rate on real traffic0.4%
10,000 samples, full pipeline
company
False-negative rate on real overeager actions17%
52 samples; 5.7% on 1,000 synthetic exfiltration cases
company

Launch and general-availability dates are from Anthropic's claude.com posts (updates dated 2026-07-10); the default-mode change was announced 2026-08-07. The error rates come from the engineering post's small samples (52 real overeager actions).

Sources

  1. www.anthropic.com/engineering/claude-code-auto-mode
  2. claude.com/blog/auto-mode
  3. claude.com/blog/auto-mode-default-in-claude-code

This record was checked against its sources on 6 October 2026. How we check

Related