AI Research Atlas

DeepSeek-V4.1-Flash

DeepSeek · 10 September 2026

Native-multimodal 552B-backbone model with a causal encoder-decoder (8B active in prefill, 16B in decode) and a KV cache of 890 bytes per token.

Adds Compressed Sparse Attention 2, FP4 KV caching and bounded sliding-window replay, cutting KV memory to about 1/4 (HBM) and 1/8 (SSD) of V4. DeepSeek says it outperforms the earlier V4-Pro on many metrics; old V4 API ids now route to it.

Date
Thursday, 10 September 2026
Lab
DeepSeek
Kind
open-weights
Access
open weights
Price
New prices from 2026-09-10 04:00 UTC; model id deepseek-flash; Flash cache-hit input $0.006 peak / $0.003 off-peak, output $1.2 / $0.6 per M

Figures

MeasureValueMeasured by
Global KV cache per token890 bytes
About 1/4 of V4-Flash; DeepSWE v1.1 74.2; Terminal-Bench 2.1 90.6
company

763B total parameters including all components per the model card. Paper arXiv 2609.19969 submitted 2026-09-17. MIT license. Parity-with-V4-Pro statements are DeepSeek's own.

Sources

  1. api-docs.deepseek.com/news/news260910
  2. huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
  3. arxiv.org/abs/2609.19969
  4. api-docs.deepseek.com/quick_start/pricing

This record was checked against its sources on 6 October 2026. How we check

Related