AI Research Atlas

DiffusionGemma technical report

Google DeepMind / Hugging Face · 31 July 2026

An experimental open-weight diffusion LM (25.2B total, 3.8B active) fine-tuned from Gemma 4 using under 10% of its training tokens; ~1,500 tokens/s on one H100.

Converts an autoregressive model into a discrete diffusion model, emitting about 20 tokens per forward pass, while keeping AR generation available with minor loss. Shows diffusion LMs can be cheaply grafted onto existing models rather than trained from scratch.

Date
Friday, 31 July 2026
Lab
Google DeepMind / Hugging Face
Kind
paper
Access
open weights

Figures

MeasureValueMeasured by
Output speed (H100)~1,500 tok/s
~20 tokens per forward pass
company

arXiv v1 submitted 2026-07-31 (43 authors, Google DeepMind and Hugging Face). Single source (arXiv abstract); no blog or release page opened.

Sources

  1. arxiv.org/abs/2608.00146

This record was checked against its sources on 6 October 2026. How we check

Related