Lumiere
Lumiere is a text-to-video diffusion model with a Space-Time U-Net that generates the whole clip in a single pass for better temporal consistency.
Avoids the keyframe-then-temporal-super-resolution cascade by down- and up-sampling in both space and time, and it builds on a pretrained text-to-image diffusion model. Also supports image-to-video, inpainting and stylized generation.
- Date
- Tuesday, 23 January 2024
- Lab
- Kind
- model
- Access
- research preview
arXiv v1 2024-01-23. Research demo only. Veo (2024-05-14) is the productized successor line, by inference.
Sources
This record was checked against its sources on 6 October 2026. How we check