AI Research Atlas

Mercury (diffusion LLMs)

Inception Labs · 17 June 2025

Mercury offers commercial-scale diffusion LLMs for code, and Mercury Coder Mini runs at 1,109 tokens/s on H100, 10x faster than speed-optimised frontier models at similar quality.

A Transformer that denoises many tokens in parallel rather than one at a time. Artificial Analysis independently measured up to 10x speed advantage; Mercury Coder ranked second on Copilot Arena quality and fastest overall. Google's Gemini Diffusion followed as a demo.

Date
Tuesday, 17 June 2025
Lab
Inception Labs
Kind
paper
Access
closed API

Figures

MeasureValueMeasured by
Throughput on H1001,109 tok/s (Mini); 737 tok/s (Small)company
Speed vs speed-optimised frontier modelsup to 10x
quality comparable per Artificial Analysis
independent

arXiv v1 2025-06-17; the product launched earlier in 2025 (date not verified here). Gemini Diffusion page reports ~1,479 tok/s and mixed quality vs Gemini 2.0 Flash-Lite (LiveCodeBench 30.9% vs 28.5%, GPQA Diamond 40.4% vs 56.5%); its announcement date was not shown on the page.

Sources

  1. arxiv.org/abs/2506.17298
  2. deepmind.google/models/gemini-diffusion/

This record was checked against its sources on 6 October 2026. How we check

Related