Gemini Diffusion
Experimental text-diffusion language model that generates blocks of tokens in parallel by refining noise; Google reports ~1,479 tokens/s with Flash-Lite-class quality.
Replaces left-to-right decoding with iterative denoising, which allows in-generation error correction and very high speed. Results against Gemini 2.0 Flash-Lite were mixed, ahead on LiveCodeBench and AIME 2025 and behind on GPQA Diamond and BIG-Bench Extra Hard. Waitlisted demo only.
- Date
- Tuesday, 20 May 2025
- Lab
- Google DeepMind
- Kind
- model
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| Sampling speed (excluding overhead) | 1,479 tokens/s plus 0.84 s overhead | company |
| LiveCodeBench v6 | 30.9% vs 28.5% for Gemini 2.0 Flash-Lite | company |
| GPQA Diamond | 40.4% vs 56.5% for Gemini 2.0 Flash-Lite | company |
Date is the I/O 2025 keynote day per press coverage. Followed by Gemma-based DiffusionGemma (2026-06-10).
Sources
- deepmind.google/models/gemini-diffusion/
- www.fortune.com/2025/05/21/gemini-diffusion-google-io-sleeper-hit-blazing-speed-ai-model-w
This record was checked against its sources on 6 October 2026. How we check