FrontierCode benchmark
Cognition releases FrontierCode, a 'mergeable code' benchmark built with 20+ open-source maintainers; the best model scores 13.4% on the hardest tier.
Scores six quality dimensions beyond passing tests (regression safety, scope, test quality). Three tiers are Extended with 150 tasks, Main with 100 and Diamond with 50. Claims 81% fewer false positives/negatives than SWE-Bench Pro. Claude Opus 4.8 tops it.
- Date
- Monday, 8 June 2026
- Lab
- Cognition
- Kind
- paper
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| Best score, Diamond tier | 13.4% (Claude Opus 4.8) Main 34.3%, Extended 51.8%; GPT-5.5 6.3% on Diamond | company |
Vendor-built benchmark; Cognition also sells and ranks models in this space. No independent replication seen.
Sources
This record was checked against its sources on 6 October 2026. How we check