AI Research Atlas

GPT-4o

OpenAI · 13 May 2024

One network handles text, audio and images end to end; GPT-4o costs half of GPT-4 Turbo and reaches free ChatGPT users.

Natively multimodal ("omni") model that processes audio in a single model instead of chaining speech-to-text, LLM and text-to-speech. A new ~200K-token tokenizer cuts non-English token use (Gujarati 4.4x fewer). Voice features were announced as rolling out later; native image generation followed in March 2025.

Date
Monday, 13 May 2024
Lab
OpenAI
Kind
model
Access
closed API
Price
$5/$15 per 1M tokens at launch; $2.50/$10 later (Wikipedia)

Figures

MeasureValueMeasured by
API price at launch$5 input / $15 output per 1M tokens
50% cheaper than GPT-4 Turbo
company
MMLU88.7 vs 86.5 for GPT-4
company-reported
company

Sky voice paused 2024-05-20 after likeness complaints; GPT-4o was removed from ChatGPT 2026-02-13 (Wikipedia). 2025 sycophancy update was a separate episode.

Sources

  1. simonwillison.net/2024/May/13/gpt-4o/
  2. en.wikipedia.org/wiki/GPT-4o
  3. developers.openai.com/api/docs/changelog

This record was checked against its sources on 6 October 2026. How we check