DeepSeek-V4-Flash-Vision-Exp
First multimodal V4 model, an experimental vision add-on to V4-Flash that matches its text ability; ships with a new Files API.
Continued training adds a vision encoder and aligner to V4-Flash. Images are billed at up to 384 tokens each at Flash prices; works over Chat Completions, Messages and Responses. Weights on Hugging Face under MIT (305B listed).
- Date
- Friday, 21 August 2026
- Lab
- DeepSeek
- Kind
- model
- Access
- open weights
- Price
- V4-Flash prices; up to 384 tokens per image
Figures
| Measure | Value | Measured by |
|---|---|---|
| ApexBench Pass@1 | 36.5 (V4-Flash 26.2) Terminal Bench 2.1 83.9 | company |
API announcement is dated 2026-08-21; Hugging Face repo creation date is 2026-08-31.
Sources
This record was checked against its sources on 6 October 2026. How we check