AI Research Atlas

Wan2.5-Preview

Alibaba (Qwen) · 24 September 2025

First Wan model with native synchronized audio (dialogue, music, effects) in one pass, up to 10-second 1080p clips; hosted only, no weights.

Natively multimodal text-image-video-audio model that generates sound and picture together, ending the silent-clip limit of Wan2.1 and 2.2. Unlike those Apache 2.0 releases, it launched as a preview API and app; Alibaba gave no open-source commitment.

Date
Wednesday, 24 September 2025
Lab
Alibaba (Qwen)
Kind
model
Access
closed API
Price
API about $0.05-$0.15 per second (The Decoder, 2025-09-25)

Figures

MeasureValueMeasured by
Clip length / resolution10 s / 1080p, 24 fps
Hosted preview; audio sync and face consistency judged imperfect by The Decoder
independent

Alibaba Cloud's Apsara post (2025-09-24) calls it a preview announced, not yet released; Model Studio lists wan2.5 t2v and i2v previews on 2025-09-23 and image models on 2025-09-24. The Decoder says Alibaba did not respond to requests for a code release.

Sources

  1. www.alibabacloud.com/blog/alibaba-cloud-unveils-strategic-roadmaps-for-the-next-generation
  2. the-decoder.com/alibabas-wan2-5-preview-lets-users-turn-photos-and-text-prompts-into-video
  3. www.alibabacloud.com/help/en/model-studio/newly-released-models

This record was checked against its sources on 6 October 2026. How we check

Related