Proceedings · Session S-116 · filed September 26, 2026

AI & Emerging Tech in R&DSession paper

GPT-6 Astra scores 62.7% on interactive reasoning benchmark

GPT-6 Astra hit 62.7% on ARC-AGI-3's neutral harness but 99.9% with OpenAI's state-preserving adapter — a gap that reshapes how R&D teams should read agent benchmarks.

By Rebecca Stone3 min read634 words

Summary

  • GPT-6 Astra scored 62.7% on ARC-AGI-3 under the standard, provider-neutral testing harness.
  • The same model scored 99.9% with an OpenAI-specific context-management adapter preserving hidden reasoning state between interactions.
  • ARC-AGI-3 is an interactive benchmark requiring agents to explore unfamiliar games, infer rules and goals, and plan actions without instructions.
GPT-6 Astra scores 62.7% on interactive reasoning benchmark, near-perfect with custom adapter
FigureGPT-6 Astra scores 62.7% on interactive reasoning benchmark, near-perfect with custom adapter — AI-generated

OpenAI's new GPT-6 Astra scored 62.7% on ARC-AGI-3 under the benchmark's standard, provider-neutral testing harness, according to results reported by Research & Development World. The same model reached 99.9% when run with an OpenAI-specific context-management adapter that preserves the model's hidden reasoning state between interactions.

The 37.2-percentage-point gap between the two configurations is the number that should hold R&D managers' attention. It is larger than the difference between frontier models on many leaderboards, and it stems not from a change in the model itself but from how the test harness connects to it.

ARC-AGI-3 differs from static benchmarks such as MMLU or GSM8K. It is an interactive evaluation: the system under test must explore unfamiliar games, infer their rules and goals, and plan effective actions without receiving explicit instructions. That design pushes against a known weakness of current language-model-based agents — the loss of accumulated reasoning context across interaction steps, since standard APIs discard the model's internal reasoning state between calls and force it to reconstruct context from the visible conversation alone.

The adapter OpenAI used addresses precisely that failure mode. By carrying the hidden reasoning state forward between interactions, it removes the reconstruction burden that the neutral harness imposes. In effect, the 62.7% figure measures GPT-6 Astra as most customers would deploy it through a conventional API, while the 99.9% figure measures it running in a configuration that preserves internal state in a way OpenAI controls.

For research groups evaluating frontier models for agentic workflows — autonomous experiment planning, multi-step data analysis, iterative literature synthesis — the result carries two practical implications.

First, benchmark scores reported under vendor-specific harnesses may not transfer to deployment conditions. A procurement decision based on the 99.9% figure assumes access to equivalent state-preserving infrastructure. Teams integrating models through standard APIs should treat the neutral-harness score, 62.7%, as the more representative estimate of out-of-the-box performance on interactive, multi-step tasks.

Second, context management is now a measurable performance lever rather than an engineering afterthought. The gap between the two scores quantifies, at least for this model-benchmark pair, how much interactive reasoning capacity goes unused when reasoning state is discarded between steps. R&D groups building agent pipelines may find that investment in context-persistence middleware — where their model stack allows it — yields larger gains than waiting for the next model generation.

The result also fits a wider pattern in AI evaluation. As benchmarks have shifted from static question-answering toward interactive, agentic assessments, the role of the harness has grown. Score comparisons across providers are meaningful only when the harness is held constant, which is what the provider-neutral configuration of ARC-AGI-3 is designed to guarantee. Vendor-tuned configurations remain useful for demonstrating a model's ceiling, but they answer a different question: what the system can do under optimal integration, not what it will do under standard conditions.

The published figures leave several questions open that lab leads will want answered before drawing portfolio conclusions. The source results do not specify how many runs produced each score, what variance accompanied them, or how the adapter's mechanism maps onto deployment options OpenAI intends to expose commercially. If the state-preserving configuration becomes a generally available API feature, the effective performance gap for customers could close substantially; if it remains internal to the vendor's own tooling, buyers should plan around the lower figure.

OpenAI has not, in the reported results, disclosed pricing or availability terms for the adapter configuration. Expect follow-up evaluations from independent labs replicating both configurations — and expect the harness question, not the headline score, to dominate the next round of frontier-model comparisons.

via arcprize.org (Original)

Filed under

  • ai-benchmarks
  • agentic-workflows
  • gpt-6
  • evaluation-methodology
  • context-management
Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Market editor covering marketplaces and e-commerce at Hypothesis Wire.

80 articles

References

  1. OpenAI shifts up to 10% of compute to safety, pauses model training
  2. Anthropic's Amodei Met Trump at White House as He Urges Slower AI Development
  3. OpenAI Reportedly Seeks $30 Billion at $1.4 Trillion Valuation
  4. Altman Labels Hugging Face Breach OpenAI's 'Worst Accident'
  5. Google's Gemini 4 Argon rejoins the frontier, but independent tests temper the claims

Next article »