AI Research Atlas

phi-1 ('Textbooks Are All You Need')

Microsoft · 20 June 2023

1.3B code model trained 4 days on 8 A100s on 'textbook quality' web plus synthetic data reaches 50.6% HumanEval.

Launched the Phi thesis that data quality can substitute for scale. The model trained on 6B filtered web tokens plus 1B GPT-3.5-generated tokens beat much larger code models on HumanEval and MBPP (55.5%).

Date
Tuesday, 20 June 2023
Lab
Microsoft
Kind
model
Access
open weights

Figures

MeasureValueMeasured by
HumanEval pass@150.6%
MBPP 55.5%; 1.3B params, 4 days on 8 A100s
company

arXiv v1 2023-06-20; Wikipedia gives a 2023-06-21 release date. Whether synthetic data leaked benchmark content was debated.

Sources

  1. arxiv.org/abs/2306.11644

This record was checked against its sources on 6 October 2026. How we check