Wan2.2 (T2V-A14B, I2V-A14B, TI2V-5B)
Open video diffusion adopts mixture-of-experts. The A14B experts split denoising by timestep, and a 5B hybrid model gives 720P 24fps on a 4090.
Wan2.2 separates the denoising process across timesteps with specialized experts, raising capacity at the same per-step compute (Alibaba cites this as new for video diffusion). Training data grew 65.6% in images and 83.2% in videos over Wan2.1. The 5B model uses a new VAE compressing 16x16x4. Apache 2.0.
- Date
- Monday, 28 July 2025
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Training data growth vs Wan2.1 | +65.6% images, +83.2% videos Per the Wan2.2 README (company) | company |
| Wan2.2-VAE compression | 16x16x4 5B model does text/image-to-video at 720P, 24fps and runs on consumer GPUs such as the 4090 | company |
README: 'official release of Wan2.2 inference code and model weights' 2025-07-28 (some weights were on Hugging Face from 2025-07-18 and 07-24). Quality claims (top among open and closed models) are company-made.
Sources
- github.com/Wan-Video/Wan2.2
- huggingface.co/Wan-AI/Wan2.2-T2V-A14B
- huggingface.co/Wan-AI/Wan2.2-TI2V-5B
This record was checked against its sources on 6 October 2026. How we check