SWE-bench Verified
OpenAI releases SWE-bench Verified, a 500-task human-validated subset of SWE-bench, which became the standard coding-agent score for 18 months.
Human annotators screened tasks for well-specified descriptions and valid tests to reduce noise and contamination in SWE-bench. In February 2026 OpenAI said it would stop quoting the score over contamination and recommended SWE-bench Pro.
- Date
- 2024-08
- Lab
- OpenAI
- Kind
- paper
- Access
- research preview
Exact day not confirmed from an opened source (training memory says 2024-08-13). Wikipedia gives August 2024 and the February 2026 retirement of the figure.
Sources
This record was checked against its sources on 6 October 2026. How we check