Qwen-Audio-3.0 and 3.1 (TTS, ASR, realtime duplex)
New Qwen-Audio line on Alibaba Cloud with TTS at 200 ms first packet, dialect-aware ASR, and a duplex realtime speech model claimed number one globally.
Qwen-Audio-3.0 TTS (Plus and Flash) arrived 2026-07-14, ASR variants 2026-07-30 and duplex realtime 2026-08-10. Audio-3.1-Realtime added a 262K-token context on 2026-09-20; Qwen-Audio-3.1-TTS-Next was shown at Apsara 2026.
- Date
- Tuesday, 14 July 2026
- Lab
- Alibaba (Qwen)
- Kind
- model
- Access
- closed API
Figures
| Measure | Value | Measured by |
|---|---|---|
| TTS-Flash first-packet latency | 200 ms Alibaba Cloud's description | company |
Dates come from Alibaba Cloud's Model Studio release list. The 'ranked #1 globally' line is the vendor's own claim and no leaderboard was verified. No open weights found.
Sources
- www.alibabacloud.com/help/en/model-studio/newly-released-models
- www.alibabagroup.com/en-US/document-2041268239081668608
This record was checked against its sources on 6 October 2026. How we check