GPT-6 Astra scores 62.7% on ARC-AGI-3 with the standard harness and 99.9% with a new provider adapter harness; Claude Opus 5 scores 30.2%, and GPT-5.6 Sol 7.8% (Greg Kamradt/ARC Prize)
claude gpt-5 openai
| Source: Techmeme | Original article
GPT-6 Astra achieved 62.7% on ARC‑AGI‑3 with the standard harness and a near‑perfect 99.9% using a new provider‑adapter harness, far surpassing Claude Opus 5’s 30.2% and GPT‑5.6 Sol’s 7.8%.
OpenAI’s flagship model GPT‑6 Astra has posted the strongest results yet on the ARC‑AGI‑3 benchmark, a high‑profile interactive test that measures an agent’s ability to explore unfamiliar games, infer rules and plan actions without explicit instructions. Using ARC Prize’s standard, provider‑neutral harness the model achieved a 62.7 % success rate at a cost of roughly $26 K per run. When the same model was run with OpenAI’s new Provider Adapter harness – which preserves opaque reasoning state between turns and compresses context for longer dialogues – its score jumped to an almost perfect 99.9 % for about $19 K.
The results place GPT‑6 Astra far ahead of its closest rivals: Anthropic’s Claude Opus 5 managed 30.2 %, while OpenAI’s own predecessor, GPT‑5.6 Sol, reached only 7.8 % under the same conditions. The stark contrast underscores how much of the performance gain stems from the provider‑specific context‑management tools rather than raw model size alone.
As we reported on 4 September, OpenAI unveiled GPT‑6 Astra after a massive training run on more than 100 000 GPUs at its Texas “Stargate” facility. These new ARC‑AGI‑3 figures confirm that the model’s improvements translate into tangible gains on demanding agentic workloads such as computer use, coding and complex reasoning.
What to watch next: OpenAI has not yet disclosed whether the Provider Adapter harness will be made available to external developers, a decision that could shape the competitive landscape for interactive AI agents. Further independent evaluations on public versions of ARC‑AGI‑3 and other benchmarks will be crucial to verify whether the near‑saturating score holds beyond the semi‑private test set. The industry will also be keen to see how rivals respond, either by adopting similar context‑management layers or by pushing raw model capabilities to close the gap.
Sources
Back to AIPULSEN