Voxtral (24B and 3B)
Voxtral, open speech-understanding models at 24B and 3B, beat Whisper large-v3 on transcription and handle 30 minutes of audio in context.
Mistral's first audio models, built on Small 3.1: transcription, Q&A and summarization from audio, function calling from voice. 32K context covers 30 minutes of transcription. API from $0.001 per minute, Apache 2.0.
- Date
- Tuesday, 15 July 2025
- Lab
- Mistral AI
- Kind
- open-weights
- Access
- open weights
- Price
- from $0.001 per minute, Jul 2025
Figures
| Measure | Value | Measured by |
|---|---|---|
| API price | from $0.001 per minute | company |
| Audio in context | 30 min transcription / 40 min understanding 32K tokens | company |
Comparative claims (vs Whisper large-v3, GPT-4o mini Transcribe, Gemini 2.5 Flash) are Mistral's own.
Sources
This record was checked against its sources on 6 October 2026. How we check