AI Research Atlas

Qwen3-ASR (0.6B, 1.7B) and Qwen3-ForcedAligner

Alibaba (Qwen) · 29 January 2026

Open speech recognition for 52 languages and dialects built on Qwen3-Omni, with a forced aligner that predicts timestamps for up to 5 minutes of speech.

Two models (1.7B best accuracy; 0.6B about 2,000x throughput at concurrency 128) with streaming and offline modes, language ID, 30 languages plus 22 Chinese dialects, and a 0.6B forced aligner covering 11 languages. Apache 2.0.

Date
Thursday, 29 January 2026
Lab
Alibaba (Qwen)
Kind
open-weights
Access
open weights

Figures

MeasureValueMeasured by
0.6B throughput at concurrency 128about 2,000x
Company-stated; accuracy-efficiency trade-off model
company
0.6B time to first token92 ms
Average TTFT; transcribes 2,000 seconds of speech per second at concurrency 128 (company)
company

README dates the release 2026-01-29 (HF repos created 2026-01-28). The hosted Qwen3-ASR-Flash API preceded it on 2025-09-08.

Sources

  1. huggingface.co/Qwen/Qwen3-ASR-1.7B
  2. github.com/QwenLM/Qwen3-ASR
  3. arxiv.org/abs/2601.21337

This record was checked against its sources on 6 October 2026. How we check

Related