One model, two harnesses

ARC Prize measured 62.7 percent for GPT-6 Astra on its own internal harness and says only that figure allows a fair comparison between vendors. OpenAI's 99.9 percent came under different conditions, on a harness that keeps reasoning chains between requests and summarises long runs. Across the same 167 game-reasoning pairs, OpenAI's harness ran about 3.66 times faster and spent 49 percent fewer tokens. The distance between the two figures is a question about where the measurement stops.[1]

Aggregate scores split the same way. Epoch AI combined more than 50 benchmarks and put Astra first at 169 points. Artificial Analysis gave the same model 61 points, level with its predecessor Sol and behind Claude Fable 5.1 at 66. One model lands in two different places under two ways of adding up.[1]

The cost curve runs backwards

On the standard ARC scaffold the cost is 49,791 dollars with no reasoning and 26,098 dollars at maximum reasoning. Over the same range the score climbs from 35.2 percent to 62.7 percent. The reason sits in how the work is done: the model solves the games in fewer moves, which means fewer model calls and fewer tokens. Move count sets the bill, and the token price comes second.[1]

One anomaly in the same table goes unexplained: the low reasoning setting lands at 17.5 percent, below the run with no reasoning at all, and ARC Prize offers no account of it. On unit price, Astra costs 2.5 times what Sol charges per unit of processed text. On coding work, its cost per task stays below half of Claude Fable 5's at identical scores. Which figure is real depends on where the boundary is drawn.[1]

The configuration the buyer cannot reach

This week both vendors produced their highest figures on a configuration the buyer cannot reproduce. OpenAI's 99.9 percent comes from its own harness, the one ARC Prize does not use for comparison. The Muse Spark 1.3 scores Meta leads with belong to the max configuration, still in a limited partner preview while safety testing finishes, while the tier developers get through Muse Code and the Meta Model API is xhigh.[1], [2]

On Meta's own figures, max leads xhigh with 1,754 Elo against 1,709 on GDPval-AA v2, and 66.9 against 57.2 on OSWorld 2.0. Terminal-Bench 2.1 breaks the order: xhigh sits at 89.2, ahead of max at 88.8. Artificial Analysis, measuring independently, put the broadly available xhigh at 61 on its intelligence index and 0.55 dollars per task. That last number is the one a buyer can purchase today.[2]

The July column comparing two Korean investments used the same distinction: the only condition that could be measured was the published one (Two gigawatts, 200 megawatts and one nonbinding term sheet, 25 July 2026). The same test works for the measurement boundary. ARC Prize says it plans to publish the numbers coming out of vendor harnesses as well, and plans to release ARC-AGI-4 in the first quarter of 2027. If both numbers appear in one place, a reader can see a model's two figures side by side.[1], [3]