Wan2.1 (T2V 1.3B and 14B, I2V 14B)
Alibaba open-sources a full video-generation suite under Apache 2.0; the 1.3B text-to-video model needs 8.19 GB VRAM, the 14B tops open rivals on company benchmarks.
Diffusion transformer with a new video VAE, trained on billions of images and videos; 14B scales with data and size per the report. Adds image-to-video, editing and bilingual (Chinese and English) on-screen text. First-last-frame (FLF2V) weights followed 2025-04-17.
- Date
- Tuesday, 25 February 2025
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| T2V-1.3B memory | 8.19 GB VRAM 5-second 480P clip on an RTX 4090 in about 4 minutes, unoptimized (company) | company |
| Technical report | arXiv 2503.20314 Submitted 2025-03-26; claims 14B beats open and commercial models on internal and external benchmarks (company) | company |
README: 'released the inference code and weights of Wan2.1' on 2025-02-25. The 'first video model to generate Chinese and English text' line is Alibaba's claim. Benchmarks are company-run; no independent leaderboard was checked.
Sources
This record was checked against its sources on 6 October 2026. How we check