Gemma 3n (E2B, E4B)
Gemma 3n is a pair of on-device multimodal open models (E2B, E4B) that use MatFormer nesting and per-layer embeddings, and E4B is the first sub-10B model above 1300 on LMArena.
Raw sizes are 5B and 8B parameters but run in about 2 GB and 3 GB of accelerator memory because per-layer embeddings sit on CPU. MatFormer trains a 2B sub-model inside the 4B model, allowing custom sizes. Takes text, image, audio and video input.
- Date
- Thursday, 26 June 2025
- Lab
- Google DeepMind
- Kind
- open-weights
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| LMArena score (E4B) | >1300 first model under 10B parameters | LMArena |
Preview weights appeared 2025-05-20 (changelog); full release 2025-06-26.
Sources
- developers.googleblog.com/en/introducing-gemma-3n-developer-guide/
- ai.google.dev/gemini-api/docs/changelog
This record was checked against its sources on 6 October 2026. How we check