DeepSeek-V4.1-Flash
Native-multimodal 552B-backbone model with a causal encoder-decoder (8B active in prefill, 16B in decode) and a KV cache of 890 bytes per token.
Adds Compressed Sparse Attention 2, FP4 KV caching and bounded sliding-window replay, cutting KV memory to about 1/4 (HBM) and 1/8 (SSD) of V4. DeepSeek says it outperforms the earlier V4-Pro on many metrics; old V4 API ids now route to it.
- Date
- Thursday, 10 September 2026
- Lab
- DeepSeek
- Kind
- open-weights
- Access
- open weights
- Price
- New prices from 2026-09-10 04:00 UTC; model id deepseek-flash; Flash cache-hit input $0.006 peak / $0.003 off-peak, output $1.2 / $0.6 per M
Figures
| Measure | Value | Measured by |
|---|---|---|
| Global KV cache per token | 890 bytes About 1/4 of V4-Flash; DeepSWE v1.1 74.2; Terminal-Bench 2.1 90.6 | company |
763B total parameters including all components per the model card. Paper arXiv 2609.19969 submitted 2026-09-17. MIT license. Parity-with-V4-Pro statements are DeepSeek's own.
Sources
- api-docs.deepseek.com/news/news260910
- huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
- arxiv.org/abs/2609.19969
- api-docs.deepseek.com/quick_start/pricing
This record was checked against its sources on 6 October 2026. How we check