AI Research Atlas

Why we no longer evaluate SWE-bench Verified

OpenAI · 23 February 2026

OpenAI says SWE-bench Verified is contaminated and flawed and recommends SWE-bench Pro, retiring its own benchmark from frontier launches.

An audit of a 27.6% subset of hard tasks found at least 59.4% had test cases that reject correct solutions, and models showed signs of training on the problems. Progress had slowed from 74.9% to 80.9% in six months.

Date
Monday, 23 February 2026
Lab
OpenAI
Kind
paper
Access
research preview

Figures

MeasureValueMeasured by
Audited tasks with flawed testsat least 59.4%
of a 27.6% subset of tasks models often failed
company

OpenAI created SWE-bench Verified (2024-08-13) and reported it for every model since; this is a self-correction of its own benchmark. Another post on 2026-07-08 questions coding evals further.

Sources

  1. openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  2. simonwillison.net/2026/Feb/19/swe-bench/

This record was checked against its sources on 6 October 2026. How we check

Related