AI Research Atlas

Qwen3-TTS (0.6B, 1.7B) with 12Hz tokenizer

Alibaba (Qwen) · 22 January 2026

Open Apache 2.0 text-to-speech models with 3-second voice cloning and voice design in 10 languages; first-packet latency as low as 97 ms.

Dual-track language model over discrete speech tokens, trained on 5M+ hours of speech; two tokenizers (a 25Hz semantic codec and a 12Hz multi-codebook low-latency codec) and streaming or non-streaming output. Natural-language control of timbre, emotion and prosody; 0.6B and 1.7B sizes.

Date
Thursday, 22 January 2026
Lab
Alibaba (Qwen)
Kind
open-weights
Access
open weights

Figures

MeasureValueMeasured by
End-to-end synthesis latencyas low as 97 ms
Streaming mode, per the model card
company
Training speech data5M+ hours, 10 languages
Per the technical report (arXiv 2601.15621, submitted 2026-01-22)
company

README dates the release 2026-01-22. The hosted Qwen3-TTS-Flash API appeared 2025-09-22 and Voice Cloning and Voice Design updates in 2025-12 (Qwen blog list). Latency is a company figure.

Sources

  1. huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
  2. github.com/QwenLM/Qwen3-TTS
  3. arxiv.org/abs/2601.15621

This record was checked against its sources on 6 October 2026. How we check

Related