Voicebox
Voicebox is a non-autoregressive flow-matching model trained on 50K+ hours of speech to infill audio, covering zero-shot TTS, noise removal, editing and style transfer.
Generalist speech model that learns in-context from audio plus text, usable across languages. Meta published the paper and demos without releasing the model.
- Date
- Friday, 23 June 2023
- Lab
- Meta (FAIR)
- Kind
- paper
- Access
- research preview
arXiv v1 2023-06-23; Meta's blog announcement was about a week earlier by memory (unverified), and the 'not released' framing is also from memory.
Sources
This record was checked against its sources on 6 October 2026. How we check