DeepSeek Coder (1.3B-33B)
DeepSeek's first public model family is a set of open code LLMs from 1.3B to 33B, trained from scratch on 2T tokens, 87% of them code.
Repository-level data ordering plus a fill-in-the-blank objective at 16K context. The 33B model led open code models at release; the 7B was reported comparable to CodeLlama-34B. Weights shipped under a custom DeepSeek Model License that allows commercial use.
- Date
- Wednesday, 1 November 2023
- Lab
- DeepSeek
- Kind
- open-weights
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| Pretraining tokens | 2T (87% code, 13% English/Chinese text) Sizes 1.3B-33B, 16K context, 86+ programming languages | company |
Date is the Hugging Face repo creation (2023-11-01 UTC); the GitHub README carries no dated announcement. Paper (arXiv 2401.14196) followed on 2024-01-25.
Sources
- github.com/deepseek-ai/DeepSeek-Coder
- arxiv.org/abs/2401.14196
- huggingface.co/api/models/deepseek-ai/deepseek-coder-33b-instruct
This record was partly confirmed: some claims could not be checked on 6 October 2026. How we check