gpt-4o-transcribe, gpt-4o-mini-transcribe and gpt-4o-mini-tts
New speech-to-text models beat Whisper on word error rate, and gpt-4o-mini-tts lets developers instruct the voice on how to speak.
Speech-to-text models built on GPT-4o with RL and extensive audio midtraining outperform Whisper across benchmarks, especially with accents and noise. The text-to-speech model is steerable by instruction (for example a "sympathetic customer service agent" tone) for the first time.
- Date
- Thursday, 20 March 2025
- Lab
- OpenAI
- Kind
- model
- Access
- closed API
Whisper-1 and these transcription models were deprecated 2026-08-26 for removal on 2027-02-26 in favor of gpt-transcribe and gpt-live-transcribe (changelog).
Sources
- openai.com/index/introducing-our-next-generation-audio-models/
- simonwillison.net/2025/Mar/20/new-openai-audio-models/
- developers.openai.com/api/docs/changelog
This record was checked against its sources on 6 October 2026. How we check