IMO 2025 gold-medal-level result (experimental reasoning model)
An experimental OpenAI reasoning model solves 5 of 6 IMO 2025 problems in natural language with no tools, 35/42 points, a gold-medal score.
Result came from general-purpose reinforcement learning and test-time compute scaling, not a math-specific system or formal verification. Proofs were graded by three former IMO medalists engaged by OpenAI, not IMO coordinators. Google DeepMind reported the same 35/42, officially graded. The model was not released.
- Date
- Saturday, 19 July 2025
- Lab
- OpenAI
- Kind
- paper
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| IMO 2025 score | 35/42 (5 of 6 problems) proofs graded by three former IMO medalists, unanimous consensus; problem 6 unsolved | mixed |
| Google DeepMind Gemini Deep Think, same exam | 35/42 officially graded; for comparison | independent |
Company-announced before the IMO-coordinated release window and graded outside the IMO process; Terence Tao and Kevin Buzzard questioned comparability to human contestants (per Wikipedia). Researcher Alexander Wei said no model with this math capability would ship for months. OpenAI page not opened; account comes from Wei's post quoted by Simon Willison.
Sources
- simonwillison.net/2025/Jul/19/openai-gold-medal-math-olympiad/
- simonwillison.net/2025/Jul/21/gemini-imo/
- github.com/aw31/openai-imo-2025-proofs
This record was checked against its sources on 6 October 2026. How we check