GPT-5.3-Codex-Spark (on Cerebras)
GPT-5.3-Codex-Spark, a smaller GPT-5.3-Codex served on Cerebras hardware, is OpenAI's first model built for near-instant real-time coding.
First milestone of the OpenAI-Cerebras partnership announced 2026-01-14; text-only with a 128K context, claimed about 1,000 tokens per second. OpenAI also cut end-to-end harness latency for all models.
- Date
- Thursday, 12 February 2026
- Lab
- OpenAI
- Kind
- model
- Access
- research preview
- Price
- ChatGPT Pro at launch
Figures
| Measure | Value | Measured by |
|---|---|---|
| Claimed generation speed | about 1,000 tokens per second company claim; per Simon Willison summary | company |
As an inference, speed as a product axis (also seen in later Ultrafast modes for GPT-5.6 and GPT-6 Astra) reflects the shift to agent loops where latency compounds.
Sources
This record was checked against its sources on 6 October 2026. How we check