DeepSeek-VL2
MoE vision-language models with 1.0B, 2.8B and 4.5B active parameters, using MLA and dynamic tiling for high-resolution images.
Moves DeepSeek-VL onto DeepSeekMoE with Multi-head Latent Attention and dynamic-tiling vision encoding; strong on OCR, charts, documents and grounding for its active-parameter size.
- Date
- Friday, 13 December 2024
- Lab
- DeepSeek
- Kind
- open-weights
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| Activated parameters | 1.0B / 2.8B / 4.5B Tiny / Small / full | company |
Sources
This record was checked against its sources on 6 October 2026. How we check