Mercury (diffusion LLMs)
Mercury offers commercial-scale diffusion LLMs for code, and Mercury Coder Mini runs at 1,109 tokens/s on H100, 10x faster than speed-optimised frontier models at similar quality.
A Transformer that denoises many tokens in parallel rather than one at a time. Artificial Analysis independently measured up to 10x speed advantage; Mercury Coder ranked second on Copilot Arena quality and fastest overall. Google's Gemini Diffusion followed as a demo.
- Date
- Tuesday, 17 June 2025
- Lab
- Inception Labs
- Kind
- paper
- Access
- closed API
Figures
| Measure | Value | Measured by |
|---|---|---|
| Throughput on H100 | 1,109 tok/s (Mini); 737 tok/s (Small) | company |
| Speed vs speed-optimised frontier models | up to 10x quality comparable per Artificial Analysis | independent |
arXiv v1 2025-06-17; the product launched earlier in 2025 (date not verified here). Gemini Diffusion page reports ~1,479 tok/s and mixed quality vs Gemini 2.0 Flash-Lite (LiveCodeBench 30.9% vs 28.5%, GPQA Diamond 40.4% vs 56.5%); its announcement date was not shown on the page.
Sources
This record was checked against its sources on 6 October 2026. How we check