How we built our multi-agent research system
An Opus 4 lead with Sonnet 4 subagents beat single-agent Opus 4 by 90.2% on Anthropic's research eval, at about 15x the tokens of chat.
Engineering account of the architecture behind Claude's Research feature. Token usage alone explained 80% of variance on BrowseComp, multi-agent systems burn about 15x the tokens of a chat (agents about 4x), so they pay off only on high-value tasks. Prompting, tool design and evaluation lessons.
- Date
- Friday, 13 June 2025
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Multi-agent vs single-agent Opus 4 on internal research eval | +90.2% Opus 4 lead agent with Sonnet 4 subagents | company |
| Token usage vs chat | about 15x (multi-agent), 4x (single agent) token usage explains 80% of BrowseComp variance | company |
Internal evaluation; the orchestrator-worker pattern is also in 'Building Effective Agents'.
Sources
This record was checked against its sources on 6 October 2026. How we check