Llama 3.2 (1B, 3B, 11B Vision, 90B Vision)
First multimodal Llamas (11B, 90B vision) and on-device 1B/3B text models with 128K context; Llama Stack introduced.
Added image understanding via adapter layers to Llama 3.1 text models, and small models tuned for phones with Arm, MediaTek and Qualcomm; Meta says 11B/90B are competitive with Claude 3 Haiku and GPT-4o-mini on image recognition.
- Date
- Wednesday, 25 September 2024
- Lab
- Meta
- Kind
- open-weights
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| Sizes | 1B, 3B text; 11B, 90B vision 1B/3B support 128K context | company |
Sources
This record was checked against its sources on 6 October 2026. How we check