Why we no longer evaluate SWE-bench Verified
OpenAI says SWE-bench Verified is contaminated and flawed and recommends SWE-bench Pro, retiring its own benchmark from frontier launches.
An audit of a 27.6% subset of hard tasks found at least 59.4% had test cases that reject correct solutions, and models showed signs of training on the problems. Progress had slowed from 74.9% to 80.9% in six months.
- Date
- Monday, 23 February 2026
- Lab
- OpenAI
- Kind
- paper
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| Audited tasks with flawed tests | at least 59.4% of a 27.6% subset of tasks models often failed | company |
OpenAI created SWE-bench Verified (2024-08-13) and reported it for every model since; this is a self-correction of its own benchmark. Another post on 2026-07-08 questions coding evals further.
Sources
- openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- simonwillison.net/2026/Feb/19/swe-bench/
This record was checked against its sources on 6 October 2026. How we check