V-JEPA
Video JEPA learns by predicting masked spatio-temporal regions in latent space from unlabeled video; frozen encoder reused across tasks.
Extended I-JEPA to video with 1.5x-6x better training/sample efficiency than generative pixel prediction, as a step toward LeCun's world-model agenda.
- Date
- Thursday, 15 February 2024
- Lab
- Meta
- Kind
- open-weights
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| Efficiency vs generative approaches | 1.5x-6x training and sample efficiency, per Meta | company |
Released under CC BY-NC (non-commercial).
Sources
This record was checked against its sources on 6 October 2026. How we check