DSpark speculative decoding and DeepSpec
Semi-autoregressive drafter with confidence-scheduled verification raises per-user generation speed 60-85% at matched throughput versus MTP-1 in V4 serving.
A three-block drafter predicts several tokens at once and verifies adaptively by load and confidence. DeepSeek released DSpark modules for V4-Pro and V4-Flash plus DeepSpec, a training and eval codebase (DSpark, DFlash, EAGLE3) with Qwen3 and Gemma-4 drafters.
- Date
- Saturday, 27 June 2026
- Lab
- DeepSeek
- Kind
- infra
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Per-user generation speedup vs MTP-1 | 60-85% At matched throughput inside DeepSeek's V4 serving system | company |
Date is the Hugging Face repo creation (2026-06-27); paper arXiv 2607.05147 (July 2026). DSpark is built into the later V4-Flash-0731 and V4-Pro-0813 checkpoints. MIT license for DeepSpec.
Sources
- arxiv.org/abs/2607.05147
- huggingface.co/deepseek-ai/DeepSeek-V4-Pro-DSpark
- github.com/deepseek-ai/DeepSpec
This record was checked against its sources on 6 October 2026. How we check