Wan-Dancer-14B
Open 14B model that turns music into minute-long, 720p/30fps dance video using keyframe planning plus local refinement.
Two-stage hierarchical design with global keyframe planning from the full track, then local temporal refinement, plus time-mapped RoPE for rhythm alignment and an optical-flow loss for motion continuity. Handles five dance genres from audio and text prompts. Apache 2.0.
- Date
- Friday, 10 July 2026
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Output length | over 1 minute at 720p/30fps Per the paper abstract (company claim); prior diffusion models usually fail beyond about 20 s | company |
Date is the Hugging Face repo creation (2026-07-10); arXiv 2607.09581 had its v3 on 2026-07-17, the first-submission date was not read.
Sources
This record was checked against its sources on 6 October 2026. How we check