GPT-4o
One network handles text, audio and images end to end; GPT-4o costs half of GPT-4 Turbo and reaches free ChatGPT users.
Natively multimodal ("omni") model that processes audio in a single model instead of chaining speech-to-text, LLM and text-to-speech. A new ~200K-token tokenizer cuts non-English token use (Gujarati 4.4x fewer). Voice features were announced as rolling out later; native image generation followed in March 2025.
- Date
- Monday, 13 May 2024
- Lab
- OpenAI
- Kind
- model
- Access
- closed API
- Price
- $5/$15 per 1M tokens at launch; $2.50/$10 later (Wikipedia)
Figures
| Measure | Value | Measured by |
|---|---|---|
| API price at launch | $5 input / $15 output per 1M tokens 50% cheaper than GPT-4 Turbo | company |
| MMLU | 88.7 vs 86.5 for GPT-4 company-reported | company |
Sky voice paused 2024-05-20 after likeness complaints; GPT-4o was removed from ChatGPT 2026-02-13 (Wikipedia). 2025 sycophancy update was a separate episode.
Sources
- simonwillison.net/2024/May/13/gpt-4o/
- en.wikipedia.org/wiki/GPT-4o
- developers.openai.com/api/docs/changelog
This record was checked against its sources on 6 October 2026. How we check