AI Research Atlas

Phi-4 (14B)

Microsoft · 12 December 2024

14B model trained on 9.8T tokens with heavy synthetic data beats GPT-4o on GPQA (56.1 vs 50.6) and MATH (80.4 vs 74.6).

Synthetic data throughout pretraining and new post-training let the student exceed its teacher (GPT-4) on STEM QA; trails GPT-4o on MMLU (84.8) and HumanEval (82.6). MIT licence, 16K context.

Date
Thursday, 12 December 2024
Lab
Microsoft
Kind
open-weights
Access
open weights

Figures

MeasureValueMeasured by
GPQA / MATH vs GPT-4o56.1 / 80.4 vs 50.6 / 74.6
model card figures; MMLU 84.8 vs 88.1
company

Initially on Azure AI Foundry; Hugging Face model card lists 2024-12-12. Technical report arXiv v1 2024-12-12.

Sources

  1. huggingface.co/microsoft/phi-4
  2. arxiv.org/abs/2412.08905

This record was checked against its sources on 6 October 2026. How we check

Related