VibeVoice (1.5B)
Open text-to-speech model for long multi-speaker conversations of up to 90 minutes and 4 speakers, in English and Chinese, MIT licence.
Combines continuous speech tokenizers with LLM plus diffusion head for podcast-length generation; ships with audible disclaimers and watermarks and explicit anti-impersonation limits.
- Date
- Tuesday, 26 August 2025
- Lab
- Microsoft
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Max length / speakers | 90 min / 4 speakers research-only guidance in model card | company |
Date from the model card's arXiv reference (2508.19205); the card says research-and-development use only and not for commercial deployment.
Sources
This record was checked against its sources on 6 October 2026. How we check