DiffusionGemma
DiffusionGemma is an Apache-2.0, 26B MoE open text-diffusion model that generates whole blocks of text at once for up to 4x faster inference on GPUs.
Adds a diffusion head to a Gemma 4 base using Gemini Diffusion research, so many tokens are denoised in parallel rather than one at a time. Labeled experimental, aimed at interactive local workflows.
- Date
- Wednesday, 10 June 2026
- Lab
- Google DeepMind
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Inference speed on dedicated GPUs | up to 4x faster vs autoregressive Gemma 4 decoding; company-reported | company |
A technical report with Hugging Face is catalogued separately (papers file, 2026-07-31).
Sources
- deepmind.google/blog/diffusiongemma-4x-faster-text-generation/
- blog.google/innovation-and-ai/technology/ai/google-ai-updates-june-2026/
This record was checked against its sources on 6 October 2026. How we check