Qwen3-TTS (0.6B, 1.7B) with 12Hz tokenizer
Open Apache 2.0 text-to-speech models with 3-second voice cloning and voice design in 10 languages; first-packet latency as low as 97 ms.
Dual-track language model over discrete speech tokens, trained on 5M+ hours of speech; two tokenizers (a 25Hz semantic codec and a 12Hz multi-codebook low-latency codec) and streaming or non-streaming output. Natural-language control of timbre, emotion and prosody; 0.6B and 1.7B sizes.
- Date
- Thursday, 22 January 2026
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| End-to-end synthesis latency | as low as 97 ms Streaming mode, per the model card | company |
| Training speech data | 5M+ hours, 10 languages Per the technical report (arXiv 2601.15621, submitted 2026-01-22) | company |
README dates the release 2026-01-22. The hosted Qwen3-TTS-Flash API appeared 2025-09-22 and Voice Cloning and Voice Design updates in 2025-12 (Qwen blog list). Latency is a company figure.
Sources
- huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
- github.com/QwenLM/Qwen3-TTS
- arxiv.org/abs/2601.15621
This record was checked against its sources on 6 October 2026. How we check