Qwen3.8-Omni-Flash
Natively multimodal agent model on the Qwen3.8-Next MoE architecture with a 1M-token context, built to plan and finish tasks, not just perceive.
Native multimodal co-training keeps text ability while transferring agentic skill to audio and video. Inherits Qwen3.8-Next's sparse MoE, extends context to 1M tokens, and targets video editing, translation and music-video generation. The report also presents the Qwen-MM-Plugins framework.
- Date
- Thursday, 17 September 2026
- Lab
- Alibaba (Qwen)
- Kind
- model
- Access
- closed API
Date 2026-09-17 from Releasebot and Alibaba Cloud's Model Studio list (realtime variant 2026-09-21); the Qwen3.8-Omni technical report (arXiv 2609.25611, 'Qwen Team') was submitted 2026-09-22 and confirms the 1M context and Qwen3.8-Next architecture. Access is closed API per Model Studio; no weights found. Benchmarks not opened.
Sources
- releasebot.io/updates/qwen
- www.versely.studio/blog/qwen-4-announced-at-apsara-2026
- www.alibabacloud.com/help/en/model-studio/newly-released-models
- github.com/QwenLM/Qwen-Live-Harness
- arxiv.org/abs/2609.25611
This record was checked against its sources on 6 October 2026. How we check