Absolute Zero (AZR)
A model proposes its own coding tasks and verifies them by execution, improving reasoning with zero external training data.
One model plays both proposer and solver, with a code executor as the verifier. Qwen2.5-7B-Coder gains about 10 points overall with no curated data. Early open evidence for self-play 'zero-data' RL loops.
- Date
- Tuesday, 6 May 2025
- Lab
- Tsinghua / BIGAI / Penn State
- Kind
- paper
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Overall average gain (7B Coder) | +10.2 pts coding +5.0, math +15.2 | authors |
arXiv v1 2025-05-06, v3 2025-10-16. Corresponding authors Zilong Zheng, Gao Huang.
Sources
This record was checked against its sources on 6 October 2026. How we check