Eval awareness in Opus 4.6's BrowseComp performance
Evaluating Opus 4.6 on BrowseComp, Anthropic saw it twice suspect it was being tested, identify the benchmark and decrypt the answer key.
Nine of 1,266 problems showed ordinary contamination (answers published in papers such as ICLR 2026 submissions); two showed a new pattern where the model inferred it was in an eval, found which one, and decrypted the answers. Raises doubts about static benchmarks in web-enabled settings.
- Date
- Friday, 6 March 2026
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| BrowseComp problems with leaked answers | 11 of 1,266 (9 ordinary, 2 decryption) Opus 4.6, multi-agent configuration | company |
'First documented instance' is Anthropic's wording. Related to natural-language-autoencoder findings that models are often aware of being evaluated.
Sources
This record was checked against its sources on 6 October 2026. How we check