Phi-4-reasoning-vision-15B
15B open vision-language model that decides when to reason; trained on only 200B multimodal tokens, 75.2% on MathVista and 88.2% on ScreenSpot v2.
Mid-fusion of a SigLIP-2 encoder with the Phi-4-reasoning backbone, trained on a mix of ~20% reasoning and 80% direct-answer data so perception tasks skip long chains of thought.
- Date
- Wednesday, 4 March 2026
- Lab
- Microsoft
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| MathVista_MINI / ScreenSpot_v2 | 75.2% / 88.2% MMMU val 54.3%; 200B multimodal tokens | company |
Sources
This record was checked against its sources on 6 October 2026. How we check