AI Research Atlas

Quantifying infrastructure noise in agentic coding evals

Anthropic · 5 February 2026

Anthropic shows resource configuration alone moves Terminal-Bench 2.0 scores by up to 6 points, often more than gaps between top models.

Strict container limits killed tasks for infrastructure reasons. Error rates fell from 5.8% at strict enforcement to 0.5% uncapped. The top and bottom resource setups differed by 6 percentage points (p < 0.01), so leaderboard margins of a few points may be infrastructure noise.

Date
Thursday, 5 February 2026
Lab
Anthropic
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Terminal-Bench 2.0 gap between best and worst resource setups6 percentage points
p < 0.01, Anthropic internal experiments
company
Infrastructure error rate5.8% strict vs 0.5% uncapped
tasks failing for container reasons
company

Anthropic measured its own harness; useful context for reading all agentic-coding leaderboard claims, including its own.

Sources

  1. www.anthropic.com/engineering/infrastructure-noise

This record was checked against its sources on 6 October 2026. How we check

Related