DeepSeek-Coder-V2
Open MoE code model (236B, 21B active; 16B Lite) claimed to beat GPT-4-Turbo, Claude 3 Opus and Gemini 1.5 Pro on code and math benchmarks.
Continues V2 pretraining on 6T more tokens; programming languages grow from 86 to 338 and context from 16K to 128K, with general-language ability preserved. First open model credibly claimed to match the top closed models on coding.
- Date
- Monday, 17 June 2024
- Lab
- DeepSeek
- Kind
- open-weights
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| Extra pretraining tokens | 6T Languages 86 to 338; context 16K to 128K | company |
Date is the arXiv submission; the API shipped deepseek-coder V2-0614 on 2024-06-14 and weights on HF repo creation 2024-06-14 (UTC). Benchmark superiority is company-claimed.
Sources
- arxiv.org/abs/2406.11931
- huggingface.co/api/models/deepseek-ai/DeepSeek-Coder-V2-Instruct
- api-docs.deepseek.com/updates
This record was checked against its sources on 6 October 2026. How we check