AI Research Atlas

FrontierCode benchmark

Cognition · 8 June 2026

Cognition releases FrontierCode, a 'mergeable code' benchmark built with 20+ open-source maintainers; the best model scores 13.4% on the hardest tier.

Scores six quality dimensions beyond passing tests (regression safety, scope, test quality). Three tiers are Extended with 150 tasks, Main with 100 and Diamond with 50. Claims 81% fewer false positives/negatives than SWE-Bench Pro. Claude Opus 4.8 tops it.

Date
Monday, 8 June 2026
Lab
Cognition
Kind
paper
Access
research preview

Figures

MeasureValueMeasured by
Best score, Diamond tier13.4% (Claude Opus 4.8)
Main 34.3%, Extended 51.8%; GPT-5.5 6.3% on Diamond
company

Vendor-built benchmark; Cognition also sells and ranks models in this space. No independent replication seen.

Sources

  1. cognition.com/blog/frontier-code

This record was checked against its sources on 6 October 2026. How we check