EmbeddingGemma 2
Google DeepMind released EmbeddingGemma 2, an open 740M parameter embedding model built on Gemma 4 that maps text, code, images, video and audio into one 768-dimensional space.
The first EmbeddingGemma handled text only. This version adds code, image, video and audio inputs in a single shared vector space, so developers no longer need to chain captioning, speech-to-text and text-embedding models. The encoders are modular, so a developer can load only the text part at 270M parameters, or add vision and audio up to 740M. Google says the text-only weights run in about 191MB of active RAM and the full model in about 567MB on a Pixel 11 Pro.
- Date
- Tuesday, 6 October 2026
- Lab
- Google DeepMind
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Parameters (full multimodal) | 740M 270M text and code only, 440M text, image and video, 570M text and audio | company |
| Embedding dimension | 768 shared across all modalities | company |
| Active RAM, text-only weights | ~191MB on a Google Pixel 11 Pro | company |
| Active RAM, full multimodal model | ~567MB on a Google Pixel 11 Pro | company |
Released under Apache 2.0. An ML Kit service on Android is promised in the coming weeks. Benchmark scores were not in the page text I read.
Sources
- deepmind.google/blog/embeddinggemma-2-an-open-lightweight-multimodal-embedding-model/
- developers.googleblog.com/embeddinggemma-2-the-developer-guide/
- developers.googleblog.com/google-ai-edge-with-embeddinggemma-2/
This record was checked against its sources on 6 October 2026. How we check