AI Research Atlas

Phi-4-reasoning and Phi-4-reasoning-plus (14B)

Microsoft · 30 April 2025

14B reasoning model SFT'd on o3-mini traces; reasoning-plus adds outcome RL. Beats DeepSeek-R1-Distill-70B on reasoning benchmarks, per Microsoft.

Showed careful curation of chain-of-thought SFT data lifts a small model near full DeepSeek-R1 on some tasks, and RL adds more; trained in about 2.5 days on 32 H100s. Phi-4-reasoning scores 62.9 AIME 2025 and 65.8 GPQA-Diamond.

Date
Wednesday, 30 April 2025
Lab
Microsoft
Kind
open-weights
Access
open weights

Figures

MeasureValueMeasured by
AIME 2025 / GPQA-Diamond (Phi-4-reasoning)62.9 / 65.8
vs DeepSeek-R1 70.4 / 73.0; o3-mini 78.0 / 77.7
company

Teacher traces came from OpenAI o3-mini, so this is distillation from a closed model.

Sources

  1. huggingface.co/microsoft/Phi-4-reasoning
  2. arxiv.org/abs/2504.21318

This record was checked against its sources on 6 October 2026. How we check

Related