AI Research Atlas

GPT-5.3-Codex-Spark (on Cerebras)

OpenAI · 12 February 2026

GPT-5.3-Codex-Spark, a smaller GPT-5.3-Codex served on Cerebras hardware, is OpenAI's first model built for near-instant real-time coding.

First milestone of the OpenAI-Cerebras partnership announced 2026-01-14; text-only with a 128K context, claimed about 1,000 tokens per second. OpenAI also cut end-to-end harness latency for all models.

Date
Thursday, 12 February 2026
Lab
OpenAI
Kind
model
Access
research preview
Price
ChatGPT Pro at launch

Figures

MeasureValueMeasured by
Claimed generation speedabout 1,000 tokens per second
company claim; per Simon Willison summary
company

As an inference, speed as a product axis (also seen in later Ultrafast modes for GPT-5.6 and GPT-6 Astra) reflects the shift to agent loops where latency compounds.

Sources

  1. openai.com/index/introducing-gpt-5-3-codex-spark/
  2. simonwillison.net/2026/Feb/12/codex-spark/

This record was checked against its sources on 6 October 2026. How we check

Related