AI Research Atlas

Voxtral (24B and 3B)

Mistral AI · 15 July 2025

Voxtral, open speech-understanding models at 24B and 3B, beat Whisper large-v3 on transcription and handle 30 minutes of audio in context.

Mistral's first audio models, built on Small 3.1: transcription, Q&A and summarization from audio, function calling from voice. 32K context covers 30 minutes of transcription. API from $0.001 per minute, Apache 2.0.

Date
Tuesday, 15 July 2025
Lab
Mistral AI
Kind
open-weights
Access
open weights
Price
from $0.001 per minute, Jul 2025

Figures

MeasureValueMeasured by
API pricefrom $0.001 per minutecompany
Audio in context30 min transcription / 40 min understanding
32K tokens
company

Comparative claims (vs Whisper large-v3, GPT-4o mini Transcribe, Gemini 2.5 Flash) are Mistral's own.

Sources

  1. mistral.ai/news/voxtral

This record was checked against its sources on 6 October 2026. How we check

Related