PaLM-E
PaLM-E feeds images, robot state and text as "multimodal sentences" to one embodied LLM; its 562B version set a state of the art on OK-VQA.
Injects continuous sensor inputs into a pretrained LLM and trains end-to-end for robot planning, VQA and captioning across embodiments, with positive transfer from joint internet-scale training. The 12B version served as one of the base models co-fine-tuned for RT-2.
- Date
- Monday, 6 March 2023
- Lab
- Kind
- model
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| Largest model size | 562B parameters PaLM-E-562B; state of the art on OK-VQA | paper |
arXiv v1 2023-03-06; Robotics at Google, TU Berlin and Google Research authors, pre-merger.
Sources
This record was checked against its sources on 6 October 2026. How we check