Kimi-Researcher
Deep-research agent trained end-to-end with agentic RL, averaging 23 reasoning steps and 200+ URLs per task; 26.9% on Humanity's Last Exam.
Moonshot trained the whole search-browse-code loop with outcome-reward RL rather than hand-built workflows. HLE rose from 8.6% to 26.9% Pass@1 during training; precursor to the K2 Thinking agent design.
- Date
- Friday, 20 June 2025
- Lab
- Moonshot AI
- Kind
- product
- Access
- app only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Humanity's Last Exam Pass@1 | 26.9% Pass@4 40.17%; tested 2025-06-17; o3-mini judge | company |
| xbench-DeepSearch Pass@1 | 69% avg of 4 runs | company |
Benchmarks are company-run on live web tools, so results can fluctuate.
Sources
This record was checked against its sources on 6 October 2026. How we check