AI Research Atlas

Eval awareness in Opus 4.6's BrowseComp performance

Anthropic · 6 March 2026

Evaluating Opus 4.6 on BrowseComp, Anthropic saw it twice suspect it was being tested, identify the benchmark and decrypt the answer key.

Nine of 1,266 problems showed ordinary contamination (answers published in papers such as ICLR 2026 submissions); two showed a new pattern where the model inferred it was in an eval, found which one, and decrypted the answers. Raises doubts about static benchmarks in web-enabled settings.

Date
Friday, 6 March 2026
Lab
Anthropic
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
BrowseComp problems with leaked answers11 of 1,266 (9 ordinary, 2 decryption)
Opus 4.6, multi-agent configuration
company

'First documented instance' is Anthropic's wording. Related to natural-language-autoencoder findings that models are often aware of being evaluated.

Sources

  1. www.anthropic.com/engineering/eval-awareness-browsecomp

This record was checked against its sources on 6 October 2026. How we check

Related