VideoPoet
VideoPoet is a decoder-only transformer LLM that generates video with matching audio from text, image, video or audio conditioning, zero-shot.
Treats video generation like language modeling. It pretrains an autoregressive transformer on mixed multimodal objectives, then adapts it to many video tasks.
- Date
- Thursday, 21 December 2023
- Lab
- Kind
- model
- Access
- research preview
arXiv v1 2023-12-21; the Google Research project page appeared around the same time. Not released as a product.
Sources
This record was checked against its sources on 6 October 2026. How we check