Emu Video and Emu Edit
Text-to-video via two diffusion stages (image, then video) and instruction-based image editing, shown as research only.
Factorised video generation into text-to-image then image-and-text-to-video, using two models instead of the five-model cascade of prior work; 512x512, 4 s at 16 fps.
- Date
- Thursday, 16 November 2023
- Lab
- Meta
- Kind
- paper
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| Human preference vs Make-A-Video (quality) | 96% 85% on text faithfulness; Meta's own human eval | company |
Not released as weights; paper and demo only.
Sources
This record was checked against its sources on 6 October 2026. How we check