AI Research Atlas

Qwen3-Omni-30B-A3B

Alibaba (Qwen) · 22 September 2025

Open natively omni-modal 30B-A3B Thinker-Talker MoE: audio latency as low as 211 ms, 19 speech-input and 10 speech-output languages, Apache 2.0.

Text-first pretraining mixed with audio and video so text and image quality do not regress. Ships as Instruct, Thinking and Captioner checkpoints. Alibaba claims leading results on 22 of 36 audio and audio-visual benchmarks (open-source best on 32) and cites about 20 million hours of audio training data.

Date
Monday, 22 September 2025
Lab
Alibaba (Qwen)
Kind
open-weights
Access
open weights

Figures

MeasureValueMeasured by
Response latency211 ms audio-only / 507 ms audio-video
Latency figures as stated by Qwen; measurement conditions not given in the blog
company
Benchmarks with overall SOTA22 of 36
Open-source SOTA on 32 of 36 audio/video benchmarks, per the model card
company
Audio training data20M hours
Per Alibaba Cloud's Apsara summary
company

Technical report arXiv 2509.17765 submitted 2025-09-22. Benchmark counts and latency are Qwen's own measurements. Hugging Face metadata lists the license as Apache 2.0.

Sources

  1. huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct
  2. qwen.ai/blog?id=qwen3-omni
  3. arxiv.org/abs/2509.17765
  4. github.com/QwenLM/Qwen3-Omni
  5. www.alibabacloud.com/blog/alibaba-cloud-unveils-strategic-roadmaps-for-the-next-generation

This record was checked against its sources on 6 October 2026. How we check

Related