AI Research Atlas

GPT-5.5 and GPT-5.5 Pro

OpenAI · 23 April 2026

GPT-5.5 reaches 82.7% on Terminal-Bench 2.0 and 78.7% on OSWorld-Verified; the API followed a day later with a 1M-token window.

Rolled out first to paid ChatGPT and Codex, with the API delayed for safeguards; Pro for harder problems. Gains concentrate in agentic coding, computer use and early scientific research with fewer tokens per task. The UK AISI measured a 71.4% cyber pass rate. Tendency to mention goblins traced to a personality reward.

Date
Thursday, 23 April 2026
Lab
OpenAI
Kind
model
Access
closed API
Price
API from 2026-04-24; 1M-token context

Figures

MeasureValueMeasured by
Terminal-Bench 2.082.7%
GPT-5.4 75.1%
company
OSWorld-Verified78.7%
GPT-5.4 75.0%
company
GDPval (wins or ties)84.9%
GPT-5.4 83.0%
company
FrontierMath Tier 435.4%
GPT-5.4 27.1%; Tiers 1-3 51.7%
company
UK AISI cyber tasks, average pass rate71.4% (+/-8.0)
comparable to Claude Mythos Preview 68.6%; figures per Wikipedia summary of the AISI report
independent

Codename "Spud" and the goblin post-mortem ("Where the goblins came from", 2026-04-29, a creature-metaphor tic traced to reward for the Nerdy personality) are from Wikipedia and OpenAI news; a case study in reward leakage across training runs. GPT-5.5 Instant followed 2026-05-05 and GPT-5.5-Cyber 2026-05-07.

Sources

  1. openai.com/index/introducing-gpt-5-5/
  2. deploymentsafety.openai.com/gpt-5-5/
  3. simonwillison.net/2026/Apr/23/gpt-5-5/
  4. simonwillison.net/2026/Apr/30/gpt-55-cyber-capabilities/
  5. developers.openai.com/api/docs/changelog

This record was checked against its sources on 6 October 2026. How we check

Related