DINOv2
Self-supervised vision backbone trained on 142M curated images, no labels or captions, rivals CLIP-style features.
Produced general-purpose frozen image features that work with simple linear heads for classification, depth and segmentation, showing label-free pretraining can match weakly-supervised text-image models.
- Date
- Monday, 17 April 2023
- Lab
- Meta
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Training data | 142M curated images curated from 1.2B source images; ViT-g/14 is 1.1B params | company |
Sources
This record was partly confirmed: some claims could not be checked on 6 October 2026. How we check