AI Research Atlas

DeepSeek-VL (1.3B, 7B)

DeepSeek · 8 March 2024

DeepSeek's first vision-language models (1.3B, 7B), built for real-world screenshots, PDFs, charts and OCR with a hybrid high-res encoder.

Hybrid vision encoder handles 1024x1024 inputs cheaply; training keeps language data in the mix throughout so text ability does not degrade while adding vision.

Date
Friday, 8 March 2024
Lab
DeepSeek
Kind
open-weights
Access
open weights (restricted license)

License type for VL weights not independently re-read; assumed DeepSeek Model License like sibling repos.

Sources

  1. arxiv.org/abs/2403.05525
  2. huggingface.co/api/models/deepseek-ai/deepseek-vl-7b-base

This record was checked against its sources on 6 October 2026. How we check