DeepSeek-R1-0528-Qwen3-8B
R1-0528's chain of thought distilled into Qwen3-8B Base, which scores 86.0% on AIME 2024, level with Qwen3-235B-thinking and 10 points above Qwen3-8B.
A cross-lab distillation in which DeepSeek post-trains Alibaba's Qwen3-8B Base on R1-0528 reasoning traces. It shows the Chinese open-weights ecosystem reusing each other's models and R1 acting as a teacher for small models.
- Date
- Thursday, 29 May 2025
- Lab
- DeepSeek
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| AIME 2024 | 86.0% Qwen3-8B 76.0%; AIME 2025 76.3% vs 67.3% | company |
Shares Qwen3-8B's architecture but uses DeepSeek's tokenizer config.
Sources
- huggingface.co/deepseek-ai/DeepSeek-R1-0528-Qwen3-8B
- huggingface.co/api/models/deepseek-ai/DeepSeek-R1-0528-Qwen3-8B
This record was checked against its sources on 6 October 2026. How we check