AI Research Atlas

RT-2

Google DeepMind · 28 July 2023

A vision-language-action model that emits robot actions as text tokens from a web-pretrained VLM, and it doubles performance on unseen tasks versus RT-1.

Co-fine-tunes a web-scale vision-language model (PaLI-X 55B or PaLM-E 12B) on robot trajectories so actions are output as tokens. Success on unseen scenarios rose from 32% (RT-1) to 62%; shows basic chain-of-thought and novel-object reasoning.

Date
Friday, 28 July 2023
Lab
Google DeepMind
Kind
model
Access
research preview

Figures

MeasureValueMeasured by
Unseen-scenario success62%
vs 32% for RT-1, 6,000+ trials
company

Describes itself as a vision-language-action (VLA) model; Gemini Robotics later builds on the same framing (inference).

Sources

  1. deepmind.google/blog/rt-2-new-model-translates-vision-and-language-into-action/
  2. arxiv.org/abs/2307.15818

This record was checked against its sources on 6 October 2026. How we check