Competitive Programming with Large Reasoning Models
General RL-trained o3 reached IOI 2024 gold without hand-built test-time strategies, beating the domain-specialised o1-ioi.
Compares o1, a hand-engineered o1-ioi pipeline, and general o3. The result was that scaling general-purpose RL beat domain-specific engineering. A rare primary-source statement of the 'bitter lesson' for reasoning models.
- Date
- Monday, 3 February 2025
- Lab
- OpenAI
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| IOI 2024 (o1-ioi, live competition) | 49th percentile gold under relaxed constraints | company |
| IOI 2024 (o3) | gold-medal level no IOI-specific engineering | company |
Lead author Ahmed El-Kishky. Company-run evaluation; not independently reproduced.
Sources
This record was checked against its sources on 6 October 2026. How we check