Quantifying infrastructure noise in agentic coding evals
Anthropic shows resource configuration alone moves Terminal-Bench 2.0 scores by up to 6 points, often more than gaps between top models.
Strict container limits killed tasks for infrastructure reasons. Error rates fell from 5.8% at strict enforcement to 0.5% uncapped. The top and bottom resource setups differed by 6 percentage points (p < 0.01), so leaderboard margins of a few points may be infrastructure noise.
- Date
- Thursday, 5 February 2026
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Terminal-Bench 2.0 gap between best and worst resource setups | 6 percentage points p < 0.01, Anthropic internal experiments | company |
| Infrastructure error rate | 5.8% strict vs 0.5% uncapped tasks failing for container reasons | company |
Anthropic measured its own harness; useful context for reading all agentic-coding leaderboard claims, including its own.
Sources
This record was checked against its sources on 6 October 2026. How we check