Phi-4 (14B)
14B model trained on 9.8T tokens with heavy synthetic data beats GPT-4o on GPQA (56.1 vs 50.6) and MATH (80.4 vs 74.6).
Synthetic data throughout pretraining and new post-training let the student exceed its teacher (GPT-4) on STEM QA; trails GPT-4o on MMLU (84.8) and HumanEval (82.6). MIT licence, 16K context.
- Date
- Thursday, 12 December 2024
- Lab
- Microsoft
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| GPQA / MATH vs GPT-4o | 56.1 / 80.4 vs 50.6 / 74.6 model card figures; MMLU 84.8 vs 88.1 | company |
Initially on Azure AI Foundry; Hugging Face model card lists 2024-12-12. Technical report arXiv v1 2024-12-12.
Sources
This record was checked against its sources on 6 October 2026. How we check