Qwen3.5-Omni (Plus, Flash and Realtime)
Proprietary omni model on the Qwen3.5 generation, with hundreds of billions of parameters, 256K context and claimed leading results on 215 audio and audio-visual subtasks.
Hybrid-attention MoE for both Thinker and Talker, trained on over 100M hours of audio-visual data; handles 10+ hours of audio and 400 seconds of 720P video. Introduces ARIA to align text and speech units for steadier streaming speech. Offered as Plus, Flash and realtime variants; no weights.
- Date
- Thursday, 26 March 2026
- Lab
- Alibaba (Qwen)
- Kind
- model
- Access
- closed API
Figures
| Measure | Value | Measured by |
|---|---|---|
| Audio and audio-visual subtasks at SOTA | 215 Qwen3.5-Omni-plus; says it beats Gemini-3.1 Pro on key audio tasks and matches it on audio-visual understanding (company) | company |
| Context / training data | 256K tokens; 100M+ hours audio-visual Hundreds of billions of parameters per the report | company |
Model Studio lists the models on 2026-03-26 (snapshot ids dated 2026-03-15); the Qwen3.5-Omni Technical Report (arXiv 2604.15804) was submitted in April (v2 2026-04-21), which matches Wikipedia's April placement. Closed API. Benchmark claims are Alibaba's.
Sources
- www.alibabacloud.com/help/en/model-studio/newly-released-models
- en.wikipedia.org/wiki/Qwen
- arxiv.org/abs/2604.15804
This record was checked against its sources on 6 October 2026. How we check