AI Research Atlas

Kimi K2 Thinking

Moonshot AI · 6 November 2025

Open-weights thinking agent that interleaves reasoning with 200-300 sequential tool calls; claims state of the art on HLE with tools (44.9%) and BrowseComp (60.2%).

K2 retrained as a thinking agent with native INT4 quantization-aware training (about 2x faster inference) and 256K context. Moonshot claimed open-model leadership over GPT-5 and Claude Sonnet 4.5 on agentic search; coding was not best.

Date
Thursday, 6 November 2025
Lab
Moonshot AI
Kind
open-weights
Access
open weights

Figures

MeasureValueMeasured by
HLE (with tools)44.9%
text-only subset; Hugging Face access blocked to avoid leakage
company
BrowseComp60.2%
human baseline 29.2%
company
SWE-bench Verified71.3%
Claude Sonnet 4.5 listed at 77.2
company

Modified MIT. The model card says HLE reaches 51.3 if Hugging Face is not blocked, which illustrates benchmark-leakage risk for browsing agents.

Sources

  1. www.kimi.com/blog/kimi-k2-thinking
  2. huggingface.co/moonshotai/Kimi-K2-Thinking
  3. simonwillison.net/tags/ai-in-china/?page=2

This record was checked against its sources on 6 October 2026. How we check

Related