Phi-4-reasoning and Phi-4-reasoning-plus (14B)
14B reasoning model SFT'd on o3-mini traces; reasoning-plus adds outcome RL. Beats DeepSeek-R1-Distill-70B on reasoning benchmarks, per Microsoft.
Showed careful curation of chain-of-thought SFT data lifts a small model near full DeepSeek-R1 on some tasks, and RL adds more; trained in about 2.5 days on 32 H100s. Phi-4-reasoning scores 62.9 AIME 2025 and 65.8 GPQA-Diamond.
- Date
- Wednesday, 30 April 2025
- Lab
- Microsoft
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| AIME 2025 / GPQA-Diamond (Phi-4-reasoning) | 62.9 / 65.8 vs DeepSeek-R1 70.4 / 73.0; o3-mini 78.0 / 77.7 | company |
Teacher traces came from OpenAI o3-mini, so this is distillation from a closed model.
Sources
This record was checked against its sources on 6 October 2026. How we check