AI Research Atlas

DeepSeek LLM 7B / 67B

DeepSeek · 29 November 2023

First general-purpose DeepSeek base and chat models (7B, 67B), trained on 2T bilingual tokens, with a scaling-law study to pick data and model allocation.

A Llama-2-style dense model with its own scaling-law analysis (the paper's 'longtermism' framing). The 67B model was reported ahead of LLaMA-2 70B on code, math and reasoning; chat variants add SFT and DPO.

Date
Wednesday, 29 November 2023
Lab
DeepSeek
Kind
open-weights
Access
open weights (restricted license)

Figures

MeasureValueMeasured by
MMLU (67B base)71.3
67B Chat scores GSM8K 84.1 and HumanEval 73.8.
company

Date is the Hugging Face repo creation (2023-11-29 UTC); the arXiv paper 2401.02954 was submitted 2024-01-05.

Sources

  1. github.com/deepseek-ai/DeepSeek-LLM
  2. arxiv.org/abs/2401.02954
  3. huggingface.co/api/models/deepseek-ai/deepseek-llm-67b-base

This record was checked against its sources on 6 October 2026. How we check