Automated Alignment Researchers
Claude agents working as automated alignment researchers closed 0.97 of a weak-to-strong supervision gap in five days, versus 0.23 for human researchers in seven.
In an Anthropic Fellows study, parallel Claude agents iterated on weak-to-strong generalization (Qwen 3-4B-Base taught by Qwen 1.5-0.5B-Chat). After human baselines took seven days to reach a 0.23 performance gap recovered, the agents reached 0.97 over 800 cumulative hours for about $18,000.
- Date
- Tuesday, 14 April 2026
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Performance gap recovered (PGR) | 0.97 vs 0.23 human baseline Agents took 5 further days, 800 hours, about $18,000 ($22 per agent-hour). Humans took 7 days. | company |
One small-model setup with a single well-defined metric, so it is evidence for automated research on measurable tasks, not on fuzzy alignment problems. Follow-up on 2026-08-28 extends it (anthropic-automated-researchers-mitigate-failures).
Sources
This record was checked against its sources on 6 October 2026. How we check