V-JEPA 2
1.2B-param video world model pretrained on 1M+ hours of video, then adapted on only 62 hours of robot data for zero-shot pick-and-place planning.
Showed a JEPA world model can plan robot actions in unseen environments (65-80% pick-and-place success per Meta) and released three physical-reasoning benchmarks (IntPhys 2, MVPBench, CausalVQA) where models sit near chance versus 85-95% for humans on IntPhys 2.
- Date
- Wednesday, 11 June 2025
- Lab
- Meta
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Robot pick-and-place success (new environments) | 65-80% zero-shot, after 62 h of action-conditioned robot data | company |
Meta's claim of being much faster than Nvidia Cosmos for planning is company-stated and was not verified from pages opened.
Sources
This record was checked against its sources on 6 October 2026. How we check