DeepSeek LLM 7B / 67B
First general-purpose DeepSeek base and chat models (7B, 67B), trained on 2T bilingual tokens, with a scaling-law study to pick data and model allocation.
A Llama-2-style dense model with its own scaling-law analysis (the paper's 'longtermism' framing). The 67B model was reported ahead of LLaMA-2 70B on code, math and reasoning; chat variants add SFT and DPO.
- Date
- Wednesday, 29 November 2023
- Lab
- DeepSeek
- Kind
- open-weights
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| MMLU (67B base) | 71.3 67B Chat scores GSM8K 84.1 and HumanEval 73.8. | company |
Date is the Hugging Face repo creation (2023-11-29 UTC); the arXiv paper 2401.02954 was submitted 2024-01-05.
Sources
- github.com/deepseek-ai/DeepSeek-LLM
- arxiv.org/abs/2401.02954
- huggingface.co/api/models/deepseek-ai/deepseek-llm-67b-base
This record was checked against its sources on 6 October 2026. How we check