GPT-5.1-Codex-Max
GPT-5.1-Codex-Max is OpenAI's first model natively trained to work across multiple context windows via compaction, handling tasks over 24 hours.
Compaction prunes history while keeping key context, so one task can span millions of tokens. At medium effort it beats GPT-5.1-Codex on SWE-bench Verified with 30% fewer thinking tokens. It stays below High cybersecurity capability and is OpenAI's most capable cyber model to date. API followed 2025-12-04.
- Date
- Wednesday, 19 November 2025
- Lab
- OpenAI
- Kind
- model
- Access
- closed API
- Price
- API from 2025-12-04
Figures
| Measure | Value | Measured by |
|---|---|---|
| SWE-bench Verified (high / xhigh effort) | 76.5% / 77.9% company-reported | company |
| Terminal-Bench 2.0 | 58.1% Gemini 3 Pro 54.2%, Sonnet 4.5 42.8% as reported | company |
Released the day after Gemini 3 Pro. System card dated 2025-11-18 on the Deployment Safety Hub; blog 2025-11-19.
Sources
- openai.com/index/gpt-5-1-codex-max/
- simonwillison.net/2025/Nov/19/gpt-51-codex-max/
- deploymentsafety.openai.com/gpt-5-1-codex-max/
- developers.openai.com/api/docs/changelog
This record was checked against its sources on 6 October 2026. How we check