DeepSeekMath-V2
Open 685B proof model that uses a trained verifier as reward. It reaches gold-level on IMO 2025 and CMO 2024 and 118/120 on Putnam 2024 with scaled test-time compute.
Moves from rewarding final answers to rewarding self-verified proofs. A verifier model scores proofs and the generator is trained against it, searching and critiquing its own proofs. Built on V3.2-Exp-Base, Apache 2.0.
- Date
- Thursday, 27 November 2025
- Lab
- DeepSeek
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Putnam 2024 | 118/120 With scaled test-time compute; gold-level IMO 2025 and CMO 2024 | company |
Scores are DeepSeek-reported on past contests; grading was by the team's own protocol, not official judges.
Sources
This record was checked against its sources on 6 October 2026. How we check