AI Research Atlas

SWE-bench Verified

OpenAI · 2024-08

OpenAI releases SWE-bench Verified, a 500-task human-validated subset of SWE-bench, which became the standard coding-agent score for 18 months.

Human annotators screened tasks for well-specified descriptions and valid tests to reduce noise and contamination in SWE-bench. In February 2026 OpenAI said it would stop quoting the score over contamination and recommended SWE-bench Pro.

Date
2024-08
Lab
OpenAI
Kind
paper
Access
research preview

Exact day not confirmed from an opened source (training memory says 2024-08-13). Wikipedia gives August 2024 and the February 2026 retirement of the figure.

Sources

  1. en.wikipedia.org/wiki/SWE-bench

This record was checked against its sources on 6 October 2026. How we check