GPT-5.5 and GPT-5.5 Pro
GPT-5.5 reaches 82.7% on Terminal-Bench 2.0 and 78.7% on OSWorld-Verified; the API followed a day later with a 1M-token window.
Rolled out first to paid ChatGPT and Codex, with the API delayed for safeguards; Pro for harder problems. Gains concentrate in agentic coding, computer use and early scientific research with fewer tokens per task. The UK AISI measured a 71.4% cyber pass rate. Tendency to mention goblins traced to a personality reward.
- Date
- Thursday, 23 April 2026
- Lab
- OpenAI
- Kind
- model
- Access
- closed API
- Price
- API from 2026-04-24; 1M-token context
Figures
| Measure | Value | Measured by |
|---|---|---|
| Terminal-Bench 2.0 | 82.7% GPT-5.4 75.1% | company |
| OSWorld-Verified | 78.7% GPT-5.4 75.0% | company |
| GDPval (wins or ties) | 84.9% GPT-5.4 83.0% | company |
| FrontierMath Tier 4 | 35.4% GPT-5.4 27.1%; Tiers 1-3 51.7% | company |
| UK AISI cyber tasks, average pass rate | 71.4% (+/-8.0) comparable to Claude Mythos Preview 68.6%; figures per Wikipedia summary of the AISI report | independent |
Codename "Spud" and the goblin post-mortem ("Where the goblins came from", 2026-04-29, a creature-metaphor tic traced to reward for the Nerdy personality) are from Wikipedia and OpenAI news; a case study in reward leakage across training runs. GPT-5.5 Instant followed 2026-05-05 and GPT-5.5-Cyber 2026-05-07.
Sources
- openai.com/index/introducing-gpt-5-5/
- deploymentsafety.openai.com/gpt-5-5/
- simonwillison.net/2026/Apr/23/gpt-5-5/
- simonwillison.net/2026/Apr/30/gpt-55-cyber-capabilities/
- developers.openai.com/api/docs/changelog
This record was checked against its sources on 6 October 2026. How we check