Wan2.5-Preview
First Wan model with native synchronized audio (dialogue, music, effects) in one pass, up to 10-second 1080p clips; hosted only, no weights.
Natively multimodal text-image-video-audio model that generates sound and picture together, ending the silent-clip limit of Wan2.1 and 2.2. Unlike those Apache 2.0 releases, it launched as a preview API and app; Alibaba gave no open-source commitment.
- Date
- Wednesday, 24 September 2025
- Lab
- Alibaba (Qwen)
- Kind
- model
- Access
- closed API
- Price
- API about $0.05-$0.15 per second (The Decoder, 2025-09-25)
Figures
| Measure | Value | Measured by |
|---|---|---|
| Clip length / resolution | 10 s / 1080p, 24 fps Hosted preview; audio sync and face consistency judged imperfect by The Decoder | independent |
Alibaba Cloud's Apsara post (2025-09-24) calls it a preview announced, not yet released; Model Studio lists wan2.5 t2v and i2v previews on 2025-09-23 and image models on 2025-09-24. The Decoder says Alibaba did not respond to requests for a code release.
Sources
- www.alibabacloud.com/blog/alibaba-cloud-unveils-strategic-roadmaps-for-the-next-generation
- the-decoder.com/alibabas-wan2-5-preview-lets-users-turn-photos-and-text-prompts-into-video
- www.alibabacloud.com/help/en/model-studio/newly-released-models
This record was checked against its sources on 6 October 2026. How we check