GPT-6 Astra gains four points after Intelligence Index overhaul
Artificial Analysis rebuilt its Intelligence Index with two new tests and 40 per cent held-out data after GPT-6 Astra's score drew skepticism. Astra gained four points over its predecessor to place second, while Claude Fable 5.1 stayed first. On the same week's Arena WebDev table, Astra Max's higher point estimate still carried the same 1–2 rank interval as its rival, showing how evaluation design can change the picture.
Artificial Intelligence··Midday
Artificial Analysis changed the index design
Artificial Analysis rebuilt its Intelligence Index after other evaluations disagreed with the picture it gave of GPT-6 Astra. The new version adds two benchmarks, AA-Briefcase and GDP.pdf, and removes GPQA-Diamond after models solved it. Held-out private test data now carries 40 per cent of the total score. The index therefore relies more heavily on questions kept away from public training data; the resulting score change comes from the revised evaluation design, rather than an update to the model itself.[1]
Astra gained four points while Claude Fable 5.1 stayed first
Claude Fable 5.1 remains first on the revised table. GPT-6 Astra moves into second with a four-point gain over its predecessor, while Meta places third. AA-Briefcase targets agentic knowledge work and GDP.pdf tests reading long PDF documents; the removed GPQA-Diamond no longer separated models effectively. Artificial Analysis also claims that Astra uses fewer tokens per task than other frontier models. That figure comes from the same evaluation provider whose scoring approach prompted the revision, so it is not an independent measurement.[1]
Arena measures a different capability and a different uncertainty
Arena's WebDev table from the same week rests on users choosing between pairs of generated frontend applications. GPT-6 Astra Max scored 1,797 from 1,199 votes, while Claude Fable 5.1 Max scored 1,762 from 2,275 votes. Astra has the higher point estimate, yet Arena gives both models a rank interval of 1–2. The Intelligence Index task mix and Arena's interface preferences answer different questions; neither establishes general superiority across coding or intelligence tasks.[2], [3]