Qwen2.5-Omni (3B, 7B)
Qwen2.5-Omni is an end-to-end Thinker-Talker model that takes text, image, audio and video and streams text and speech. The 7B is open under Apache 2.0.
Introduces TMRoPE to time-align video with audio and a Thinker-Talker split so text and speech generate together in real time. Available on Hugging Face, ModelScope, DashScope and GitHub with a paper; the 3B size followed under the Qwen Research license.
- Date
- Thursday, 27 March 2025
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Blog dated 2025-03-27 in the feed (HF repo 2025-03-22). Licenses per Wikipedia. Benchmarks not reproduced here.
Sources
This record was checked against its sources on 6 October 2026. How we check