LLaDA (Large Language Diffusion Models)
An 8B masked-diffusion language model trained from scratch matches Llama 3 8B in-context learning and beats GPT-4o on a reversal-poem task.
Shows that bidirectional masked diffusion with iterative unmasking scales, fine-tunes to instruction following and avoids the reversal curse, so autoregression is not the only route to LLM capability. Triggered a wave of diffusion LMs (Dream, Mercury, Gemini Diffusion, LLaDA 2).
- Date
- Friday, 14 February 2025
- Lab
- Renmin University of China et al.
- Kind
- paper
- Access
- open weights
arXiv v1 2025-02-14. Authors incl. Shen Nie, Fengqi Zhu, Chongxuan Li, Ji-Rong Wen; institutions were not shown on the abstract page I opened, so the lab field is a best-effort label.
Sources
This record was checked against its sources on 6 October 2026. How we check