RT-2
A vision-language-action model that emits robot actions as text tokens from a web-pretrained VLM, and it doubles performance on unseen tasks versus RT-1.
Co-fine-tunes a web-scale vision-language model (PaLI-X 55B or PaLM-E 12B) on robot trajectories so actions are output as tokens. Success on unseen scenarios rose from 32% (RT-1) to 62%; shows basic chain-of-thought and novel-object reasoning.
- Date
- Friday, 28 July 2023
- Lab
- Google DeepMind
- Kind
- model
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| Unseen-scenario success | 62% vs 32% for RT-1, 6,000+ trials | company |
Describes itself as a vision-language-action (VLA) model; Gemini Robotics later builds on the same framing (inference).
Sources
- deepmind.google/blog/rt-2-new-model-translates-vision-and-language-into-action/
- arxiv.org/abs/2307.15818
This record was checked against its sources on 6 October 2026. How we check