Genie (generative interactive environments)
11B-parameter world model trained on unlabeled internet videos that turns an image into a playable 2D environment.
Learns a latent action space without any action labels, using a video tokenizer, latent action model and autoregressive dynamics model. Generates frame-by-frame controllable environments from text, images or sketches; learned actions let agents imitate behaviour from unseen videos.
- Date
- Friday, 23 February 2024
- Lab
- Google DeepMind
- Kind
- paper
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| Parameters | 11B trained on unlabeled internet video | company |
arXiv v1 2024-02-23. Paper only; no public model release at the time. Genie 2 (2024-12-04) and Genie 3 (2025-08-05) followed.
Sources
This record was checked against its sources on 6 October 2026. How we check