Gemma 4 12B
Gemma 4 12B: open 12B model with no separate vision or audio encoders, running in 16GB of memory, with native audio input and MTP drafters.
Vision and audio inputs flow straight into the language model, a "unified, encoder-free" design. Performance nears the 26B MoE at under half the total memory. Apache 2.0, first mid-sized Gemma with native audio input. Gemma 4 had passed 150 million downloads.
- Date
- Wednesday, 3 June 2026
- Lab
- Google DeepMind
- Kind
- open-weights
- Access
- open weights
Page dated 2026-06-03 (the DeepMind RSS feed lists 2026-06-09).
Sources
- deepmind.google/blog/introducing-gemma-4-12b-a-unified-encoder-free-multimodal-model/
- blog.google/innovation-and-ai/technology/ai/google-ai-updates-june-2026/
This record was checked against its sources on 6 October 2026. How we check