Pixtral 12B
Pixtral 12B is Mistral's first multimodal model, built on a Nemo-based 12B decoder with a new 400M vision encoder. It has 128K context and is under Apache 2.0.
Natively multimodal. Images enter as 16x16-patch tokens at any size or aspect ratio, and text-only performance is kept. Scores 52.5% on MMMU. Weights are on Hugging Face.
- Date
- Tuesday, 17 September 2024
- Lab
- Mistral AI
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| MMMU | 52.5% | company |
Mistral's announcement post is dated 2024-09-17; I recall the weights appearing a few days earlier but did not confirm that in a source, so the blog date is used.
Sources
This record was checked against its sources on 6 October 2026. How we check