AI Research Atlas

Competitive Programming with Large Reasoning Models

OpenAI · 3 February 2025

General RL-trained o3 reached IOI 2024 gold without hand-built test-time strategies, beating the domain-specialised o1-ioi.

Compares o1, a hand-engineered o1-ioi pipeline, and general o3. The result was that scaling general-purpose RL beat domain-specific engineering. A rare primary-source statement of the 'bitter lesson' for reasoning models.

Date
Monday, 3 February 2025
Lab
OpenAI
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
IOI 2024 (o1-ioi, live competition)49th percentile
gold under relaxed constraints
company
IOI 2024 (o3)gold-medal level
no IOI-specific engineering
company

Lead author Ahmed El-Kishky. Company-run evaluation; not independently reproduced.

Sources

  1. arxiv.org/abs/2502.06807

This record was checked against its sources on 6 October 2026. How we check