DeepSeek-VL (1.3B, 7B)
DeepSeek's first vision-language models (1.3B, 7B), built for real-world screenshots, PDFs, charts and OCR with a hybrid high-res encoder.
Hybrid vision encoder handles 1024x1024 inputs cheaply; training keeps language data in the mix throughout so text ability does not degrade while adding vision.
- Date
- Friday, 8 March 2024
- Lab
- DeepSeek
- Kind
- open-weights
- Access
- open weights (restricted license)
License type for VL weights not independently re-read; assumed DeepSeek Model License like sibling repos.
Sources
This record was checked against its sources on 6 October 2026. How we check