VALL-E
VALL-E treats text-to-speech as language modelling over neural-codec tokens and clones a voice from a 3-second sample, trained on 60K hours.
Replaced continuous-signal TTS with a GPT-style codec language model, and in-context learning gave zero-shot voice cloning that preserves emotion and acoustic environment. Not released publicly, it still defined the architecture many later voice-cloning products copied.
- Date
- Thursday, 5 January 2023
- Lab
- Microsoft Research
- Kind
- paper
- Access
- research preview
Microsoft did not release weights (misuse concerns, per its project page and press; not re-opened here).
Sources
This record was checked against its sources on 6 October 2026. How we check