A fixed list and an opponent that answers back

GPT-6 Astra's system card gives two resistance figures. On a fixed dataset the model refuses harmful requests at rates between 91.5 percent and 98.3 percent across biology, violence and cybersecurity. Under adaptive jailbreak attempts that run several rounds, resistance falls to roughly 67 percent, where earlier GPT-5 models sat just under 50 percent.[1]

The opponent, rather than the subject matter, separates the two. A fixed dataset asks the same questions however the model answers; the adaptive figure comes from an attacker who reads each refusal and rewrites the request. Under that pressure at least one problematic answer appears roughly one try in three. Another reading stays available: the adaptive setting may simply be harder in a way that maps poorly onto real misuse, since an attacker with unlimited retries sits far from the median user. Either way, one label cannot carry both numbers.[1]

Where the indirect number lands

The direct-injection result and the indirect one move apart. Against prompt injections placed straight into the request, Astra defends at 99.99 percent, which OpenAI credits to its GPT-Red method. On Gray Swan's IPI Arena, where the instruction hides inside content the model has to read, the failure rate is 8.5 percent across 1,810 curated attacks with 15 attempts per scenario. GPT-5.6 Sol fails 27 percent of the time there, and Claude Opus 5 fails 4.8 percent.[1]

That distinction decides who the number applies to. Astra reached general availability across GitHub Copilot's paid tiers and the editors GitHub supports, including the coding agent and the CLI, and those surfaces exist to read repository content. The indirect figure is the one that describes that setting, and the 99.99 percent direct-injection defence does not extend to it. GitHub's description of the model rests on its own internal testing, so the safety measurement and the deployment claim come from different kinds of evidence.[1], [2]

What would settle this?

Two days ago, asking what the Cyber Weapon Index measures when the vulnerabilities are planted, I argued that a security score can largely describe its own test bed: with planted vulnerabilities every model landed near the ceiling, while against real bugs all nine frontier API models scored zero. Astra's pair of jailbreak figures shows a related and separate failure. Here the test bed is realistic enough; what changes is whether the attacker adapts. A benchmark can hold its content fixed and still lose the property it was built to measure, because resistance to a static list and resistance to a responding adversary are separate quantities.[3], [1]

The next useful evidence would be an adaptive evaluation run by someone other than the model's developer, reporting the attacker's budget — attempts per scenario, rounds allowed, whether the attacker sees the refusal text — alongside the score. If an independent group publishes an adaptive jailbreak result for Astra with that budget stated by 31 December 2026, the resistance figure becomes comparable across models; without it, a static refusal rate and an adaptive one will keep being read as one property.[1]