How two API settings tripled our ARC-AGI-3 scores
OpenAI shows retained reasoning and compaction tripled GPT-5.6 Sol's ARC-AGI-3 score and cut output tokens 6x, so harness choices dominate agent benchmarks.
GPT-5.6 Sol scored 7.8% and GPT-5.5 0.4% on ARC-AGI-3 until the two settings used in ChatGPT and Codex were enabled. Argues benchmarks measure API settings, harness and prompting, not just the model; GPT-6 Astra later reported 99.9% with that harness while ARC Prize reported 62.7% in its default harness.
- Date
- Wednesday, 29 July 2026
- Lab
- OpenAI
- Kind
- paper
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| GPT-5.6 Sol on ARC-AGI-3 before settings | 7.8% GPT-5.5 0.4% | company |
| GPT-6 Astra on ARC-AGI-3, default ARC-AGI harness | 62.7% at about $26K vs 99.9% at about $19K with OpenAI's harness; ARC Prize figures per Simon Willison | independent |
On benchmark integrity, the 99.9% headline depends on a vendor-specific "Provider Adapter harness" that preserves opaque reasoning state.
Sources
- openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/
- simonwillison.net/2026/Sep/3/gpt6-astra/
This record was not yet confirmed by its sources on 6 October 2026. How we check