AI Research Atlas

Genie (generative interactive environments)

Google DeepMind · 23 February 2024

11B-parameter world model trained on unlabeled internet videos that turns an image into a playable 2D environment.

Learns a latent action space without any action labels, using a video tokenizer, latent action model and autoregressive dynamics model. Generates frame-by-frame controllable environments from text, images or sketches; learned actions let agents imitate behaviour from unseen videos.

Date
Friday, 23 February 2024
Lab
Google DeepMind
Kind
paper
Access
research preview

Figures

MeasureValueMeasured by
Parameters11B
trained on unlabeled internet video
company

arXiv v1 2024-02-23. Paper only; no public model release at the time. Genie 2 (2024-12-04) and Genie 3 (2025-08-05) followed.

Sources

  1. arxiv.org/abs/2402.15391

This record was checked against its sources on 6 October 2026. How we check