Voxtral TTS
Voxtral TTS is a 4B open-weights text-to-speech model with 3-second voice cloning, 9 languages and about 70ms model latency.
Mistral's first speech synthesis model. Zero-shot cross-lingual voice adaptation, 9.7x real-time factor, priced at $0.016 per 1,000 characters. Weights are CC BY-NC 4.0, so non-commercial. Mistral reports human evals favoring it over ElevenLabs Flash v2.5.
- Date
- Monday, 23 March 2026
- Lab
- Mistral AI
- Kind
- open-weights
- Access
- open weights (restricted license)
- Price
- $0.016 per 1,000 characters, Mar 2026
Figures
| Measure | Value | Measured by |
|---|---|---|
| Model latency | ~70ms typical inputs; real-time factor ~9.7x | company |
| Voice-customization preference | 68.4% win rate human eval vs competitor per Mistral | company |
Win rate comes from Mistral's own human evaluation.
Sources
This record was checked against its sources on 6 October 2026. How we check