Kimi K2 Thinking
Open-weights thinking agent that interleaves reasoning with 200-300 sequential tool calls; claims state of the art on HLE with tools (44.9%) and BrowseComp (60.2%).
K2 retrained as a thinking agent with native INT4 quantization-aware training (about 2x faster inference) and 256K context. Moonshot claimed open-model leadership over GPT-5 and Claude Sonnet 4.5 on agentic search; coding was not best.
- Date
- Thursday, 6 November 2025
- Lab
- Moonshot AI
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| HLE (with tools) | 44.9% text-only subset; Hugging Face access blocked to avoid leakage | company |
| BrowseComp | 60.2% human baseline 29.2% | company |
| SWE-bench Verified | 71.3% Claude Sonnet 4.5 listed at 77.2 | company |
Modified MIT. The model card says HLE reaches 51.3 if Hugging Face is not blocked, which illustrates benchmark-leakage risk for browsing agents.
Sources
- www.kimi.com/blog/kimi-k2-thinking
- huggingface.co/moonshotai/Kimi-K2-Thinking
- simonwillison.net/tags/ai-in-china/?page=2
This record was checked against its sources on 6 October 2026. How we check