Qwen3-Omni-30B-A3B
Open natively omni-modal 30B-A3B Thinker-Talker MoE: audio latency as low as 211 ms, 19 speech-input and 10 speech-output languages, Apache 2.0.
Text-first pretraining mixed with audio and video so text and image quality do not regress. Ships as Instruct, Thinking and Captioner checkpoints. Alibaba claims leading results on 22 of 36 audio and audio-visual benchmarks (open-source best on 32) and cites about 20 million hours of audio training data.
- Date
- Monday, 22 September 2025
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Response latency | 211 ms audio-only / 507 ms audio-video Latency figures as stated by Qwen; measurement conditions not given in the blog | company |
| Benchmarks with overall SOTA | 22 of 36 Open-source SOTA on 32 of 36 audio/video benchmarks, per the model card | company |
| Audio training data | 20M hours Per Alibaba Cloud's Apsara summary | company |
Technical report arXiv 2509.17765 submitted 2025-09-22. Benchmark counts and latency are Qwen's own measurements. Hugging Face metadata lists the license as Apache 2.0.
Sources
- huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct
- qwen.ai/blog?id=qwen3-omni
- arxiv.org/abs/2509.17765
- github.com/QwenLM/Qwen3-Omni
- www.alibabacloud.com/blog/alibaba-cloud-unveils-strategic-roadmaps-for-the-next-generation
This record was checked against its sources on 6 October 2026. How we check