DeepSeek-Prover-V1.5
Lean 4 prover that adds RL from proof-assistant feedback and an exploration-driven tree search (RMaxTS), reaching 63.5% on miniF2F.
SFT on formal data, RL using Lean verification as the reward, then RMaxTS, a Monte-Carlo tree search with intrinsic-reward exploration to diversify proof paths. Roughly 11 points over V1 on miniF2F.
- Date
- Thursday, 15 August 2024
- Lab
- DeepSeek
- Kind
- open-weights
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| miniF2F-test | 63.5% ProofNet 25.3% | company |
Sources
This record was checked against its sources on 6 October 2026. How we check