GPT-6 Astra scored 62.7% on the neutral test, not 99.9%
Sasha / Models and Research desk
The AGI era now carries an asterisk. On the one benchmark harness every model runs the same way, GPT-6 Astra scored 62.7%, not the 99.9% in OpenAI’s launch material.
What the standard harness showed
ARC Prize ran Astra on ARC-AGI-3 with its standard setup, the provider-neutral configuration used for every lab’s models, at a compute cost of about $26,000. The result was 62.7%.
The 99.9% figure is real, but it requires OpenAI’s own context-management adapter, which preserves the model’s hidden reasoning state between calls. On the shared rig, that state gets wiped between calls, as it does for every other model.
Strong, with a caveat written by the referee
The neutral run still contained a milestone: Astra used fewer actions than the median tested human on 96% of levels, and ARC Prize described the result as a step change in frontier capability. The same organization added that its environments are bounded and deterministic, so saturating the benchmark is not proof of AGI.
Independent scoreboards disagree with the launch framing too. Artificial Analysis rates Astra 61.2 against Claude Fable 5.1’s 65.7.
The gap between 62.7 and 99.9 is not a rounding dispute; it is the difference between a shared test and a vendor-optimized one. Which number gets cited will depend on who is doing the citing.
Sources
ANOTHER News is published by ANOTHER, an AI-native content agency. Daily coverage also runs on Instagram.