DeepSeek-V4 (CSA/HCA attention, mHC, Muon)
1.6T-parameter V4-Pro (49B active) and 284B V4-Flash (13B active) with 1M-token context; Pro uses 27% of V3.2's inference FLOPs and 10% of its KV cache.
Hybrid compressed-sparse and heavily-compressed attention for million-token context, mHC residuals and the Muon optimiser, trained on over 32T tokens. Context length of 1M is the default across DeepSeek's services; open weights.
- Date
- Friday, 24 April 2026
- Lab
- DeepSeek
- Kind
- paper
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| V4-Pro inference FLOPs vs V3.2 at 1M tokens | 27% KV cache 10%; company-reported | company |
| Pretraining tokens | >32T | company |
Preview release 2026-04-24 per DeepSeek API docs; arXiv v1 submission 2026-04-26 per the submission-history line (the arXiv id prefix 2606 does not match that date, flagged for a later check).
Sources
This record was checked against its sources on 6 October 2026. How we check