Qwen3-ASR (0.6B, 1.7B) and Qwen3-ForcedAligner
Open speech recognition for 52 languages and dialects built on Qwen3-Omni, with a forced aligner that predicts timestamps for up to 5 minutes of speech.
Two models (1.7B best accuracy; 0.6B about 2,000x throughput at concurrency 128) with streaming and offline modes, language ID, 30 languages plus 22 Chinese dialects, and a 0.6B forced aligner covering 11 languages. Apache 2.0.
- Date
- Thursday, 29 January 2026
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| 0.6B throughput at concurrency 128 | about 2,000x Company-stated; accuracy-efficiency trade-off model | company |
| 0.6B time to first token | 92 ms Average TTFT; transcribes 2,000 seconds of speech per second at concurrency 128 (company) | company |
README dates the release 2026-01-29 (HF repos created 2026-01-28). The hosted Qwen3-ASR-Flash API preceded it on 2025-09-08.
Sources
This record was checked against its sources on 6 October 2026. How we check