Kimi k1.5
Multimodal long-CoT RL model matching o1 on math, code and vision benchmarks, with a paper detailing the recipe, released the same day as DeepSeek-R1.
Open-recipe account of scaling RL on LLMs, with 128K context scaling and online mirror descent policy optimization, and no MCTS, value functions or process reward models. It also uses 'long2short' distillation to short-CoT models. Weights not released.
- Date
- Monday, 20 January 2025
- Lab
- Moonshot AI
- Kind
- model
- Access
- app only
Figures
| Measure | Value | Measured by |
|---|---|---|
| AIME 2024 (long-CoT) | 77.5 arXiv abstract | company |
| MATH-500 (long-CoT) | 96.2 | company |
| Codeforces (long-CoT) | 94th percentile | company |
| AIME (short-CoT) | 60.8 vs GPT-4o and Claude 3.5 Sonnet baselines | company |
Release date 2025-01-20 per Moonshot's GitHub/blog and Wikipedia; arXiv v1 is 2025-01-22. Benchmarks are company-reported. Overshadowed in coverage by DeepSeek-R1 on the same day.
Sources
This record was checked against its sources on 6 October 2026. How we check