DeepSeek-V4-Pro and V4-Flash (preview)
1.6T-parameter MoE (49B active) plus 284B Flash, 1M-token context; at 1M tokens uses 27% of V3.2's FLOPs and 10% of its KV cache.
Hybrid attention (Compressed Sparse plus Heavily Compressed), manifold-constrained hyper-connections and the Muon optimizer, pretrained on 32T+ tokens. Open weights under MIT in FP4/FP8; API ids deepseek-v4-pro and -flash, with the old chat/reasoner ids retired 2026-07-24.
- Date
- Friday, 24 April 2026
- Lab
- DeepSeek
- Kind
- open-weights
- Access
- open weights
- Price
- See deepseek-v4-pro-0813 and deepseek-v4-1-flash for later pricing
Figures
| Measure | Value | Measured by |
|---|---|---|
| SWE-bench Verified (V4-Pro-Max) | 80.6 Flash-Max 79.0; GPQA Diamond 90.1 vs 88.1; Codeforces 3206 vs 3052; BrowseComp 83.4 vs 73.2 | company |
| 1M-token inference cost vs V3.2 | 27% of FLOPs, 10% of KV cache Per tech report abstract | company |
No R2 shipped; V4 is the flagship after V3.2. Announcement 2026-04-24; HF repos created 2026-04-22. All benchmark numbers (SWE Verified 80.6, GPQA Diamond 90.1, Codeforces 3206, BrowseComp 83.4 for V4-Pro-Max) match the model card and are DeepSeek-reported. The tech report's arXiv page (ID 2606.19348, 'DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence', 319 authors) shows v1 submitted 2026-04-26 although the ID prefix reads as June 2026; the discrepancy is unresolved. Preview ran text-only; Flash has 13B active parameters.
Sources
- api-docs.deepseek.com/news/news260424
- huggingface.co/deepseek-ai/DeepSeek-V4-Pro
- arxiv.org/abs/2606.19348
- api-docs.deepseek.com/updates
- huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/README.md
This record was checked against its sources on 6 October 2026. How we check