Multi-token Prediction
Training with n heads predicting the next n tokens lifts HumanEval 12% and MBPP 17% at 13B and speeds decoding up to 3x.
Auxiliary future-token heads on a shared trunk improve sample efficiency, especially on code, and double as speculative-decoding drafters. DeepSeek-V3 adopted a multi-token prediction objective, bringing the idea into a frontier-scale model.
- Date
- Tuesday, 30 April 2024
- Lab
- Meta FAIR
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| HumanEval / MBPP gain (13B) | +12% / +17% vs next-token baseline; 4-token heads | authors |
| Inference speedup | up to 3x self-speculative decoding | authors |
Authors are Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Roziere, David Lopez-Paz and Gabriel Synnaeve.
Sources
This record was checked against its sources on 6 October 2026. How we check