o3 (announcement)
o3 scores 75.7% (high-efficiency) and 87.5% (low-efficiency) on ARC-AGI-1 semi-private tasks; announced on day 12 of OpenAI's December event, not released.
Scaled RL reasoning with far more test-time compute, and the 87.5% run used about 172x the compute of the 75.7% run. ARC Prize verified results. Name skipped o2. o3-mini shipped 2025-01-31; o3 itself on 2025-04-16.
- Date
- Friday, 20 December 2024
- Lab
- OpenAI
- Kind
- model
- Access
- research preview
- Price
- not available at announcement
Figures
| Measure | Value | Measured by |
|---|---|---|
| ARC-AGI-1 Semi-Private (100 tasks), high-efficiency | 75.7% $2,680 total cost for 100 tasks | independent |
| ARC-AGI-1 Semi-Private, low-efficiency | 87.5% $456,000 total cost for 100 tasks | independent |
| GPQA Diamond / SWE-bench Verified / Codeforces Elo | 87.7% / 71.7% / 2727 announcement figures for o3 | company |
Scores published by the ARC Prize Foundation on its semi-private and public sets; only the high-efficiency run stayed within the $10k public-leaderboard compute budget. The deliberative alignment paper was published the same day (covered in the papers dataset).
Sources
- arcprize.org/blog/oai-o3-pub-breakthrough
- simonwillison.net/2024/Dec/20/openai-o3-breakthrough/
- en.wikipedia.org/wiki/OpenAI_o3
This record was checked against its sources on 6 October 2026. How we check