AI Research Atlas

gpt-4o-transcribe, gpt-4o-mini-transcribe and gpt-4o-mini-tts

OpenAI · 20 March 2025

New speech-to-text models beat Whisper on word error rate, and gpt-4o-mini-tts lets developers instruct the voice on how to speak.

Speech-to-text models built on GPT-4o with RL and extensive audio midtraining outperform Whisper across benchmarks, especially with accents and noise. The text-to-speech model is steerable by instruction (for example a "sympathetic customer service agent" tone) for the first time.

Date
Thursday, 20 March 2025
Lab
OpenAI
Kind
model
Access
closed API

Whisper-1 and these transcription models were deprecated 2026-08-26 for removal on 2027-02-26 in favor of gpt-transcribe and gpt-live-transcribe (changelog).

Sources

  1. openai.com/index/introducing-our-next-generation-audio-models/
  2. simonwillison.net/2025/Mar/20/new-openai-audio-models/
  3. developers.openai.com/api/docs/changelog

This record was checked against its sources on 6 October 2026. How we check

Related