GPT-6 Astra scored 62.7 percent on ARC Prize's own test after OpenAI touted 99.9 percent
ARC Prize says only the 62.7 percent its Standard harness measured allows a fair comparison between vendors, while OpenAI's headline 99.9 percent came from a Provider Adapter setup that keeps Astra's reasoning between moves. Independent trackers Artificial Analysis and Epoch AI placed Astra anywhere from tied with its predecessor to clearly ahead, depending on which index is read.
Artificial Intelligence··Evening
Two harnesses, two very different scores
ARC Prize measured GPT-6 Astra at 62.7 percent on its own Standard harness, which strips the model's memory after every move so scores from different companies can be compared fairly. OpenAI's headline 99.9 percent instead came from a Provider Adapter setup that keeps Astra's reasoning state between moves and compacts long conversations, and OpenAI's launch-day chart paired that adapted score against rivals measured only on the plain Standard harness. Across the same 167 game-reasoning pairs, the Provider Adapter ran about 3.66 times faster and spent 49 percent fewer tokens than the Standard version, showing how much the setup itself moves the numbers.[1], [2]
Cost falls as the score climbs, but one setting runs backward
On the Standard scaffold, cost fell from 49,791 dollars with no reasoning to 26,098 dollars at maximum reasoning while the score climbed from 35.2 percent to 62.7 percent, so more reasoning effort came with both a cheaper and a stronger run. The pattern breaks at the low reasoning setting, which landed at just 17.5 percent, below the run with no reasoning at all, and ARC Prize offers no explanation for that dip. The organisation does not treat any of the results as evidence of general artificial intelligence and plans to release ARC-AGI-4 in the first quarter of 2027.[1]
Independent trackers land in different places
Independent benchmark trackers read Astra differently depending on which index they use: Epoch AI put the model in first place at 169 points, while Artificial Analysis rated it 61, level with its predecessor and behind Claude Fable 5.1 at 66. Artificial Analysis's own split indexes show the same pattern inside one tracker, placing Astra at 61.2 points against 60.9 for GPT-5.6 Sol on its broad Intelligence Index, a near-tie, but at 67.0 against 65.1 on its narrower Coding Agent Index, a clearer lead. The gap says less about what the model does than about which harness, adapter and index measured it.[1], [2]