AI Research Atlas

How two API settings tripled our ARC-AGI-3 scores

OpenAI · 29 July 2026

OpenAI shows retained reasoning and compaction tripled GPT-5.6 Sol's ARC-AGI-3 score and cut output tokens 6x, so harness choices dominate agent benchmarks.

GPT-5.6 Sol scored 7.8% and GPT-5.5 0.4% on ARC-AGI-3 until the two settings used in ChatGPT and Codex were enabled. Argues benchmarks measure API settings, harness and prompting, not just the model; GPT-6 Astra later reported 99.9% with that harness while ARC Prize reported 62.7% in its default harness.

Date
Wednesday, 29 July 2026
Lab
OpenAI
Kind
paper
Access
research preview

Figures

MeasureValueMeasured by
GPT-5.6 Sol on ARC-AGI-3 before settings7.8%
GPT-5.5 0.4%
company
GPT-6 Astra on ARC-AGI-3, default ARC-AGI harness62.7% at about $26K
vs 99.9% at about $19K with OpenAI's harness; ARC Prize figures per Simon Willison
independent

On benchmark integrity, the 99.9% headline depends on a vendor-specific "Provider Adapter harness" that preserves opaque reasoning state.

Sources

  1. openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/
  2. simonwillison.net/2026/Sep/3/gpt6-astra/

This record was not yet confirmed by its sources on 6 October 2026. How we check

Related