Eigen RadarAI
Analysis

Booz Allen's Cyber Weapon Index gap narrows once Claude Sonnet gets a cheap attack harness

Booz Allen's 18-model Cyber Weapon Index found only Claude Mythos completing a full intrusion chain unaided, scoring 80 against Claude Sonnet 5's 13. TNW's review of the same release reports that pairing Sonnet 5 with an inexpensive attack harness closed that 67-point gap, and flags a separate finding: one model refused a task for lacking credentials while its cyber-tuned sibling completed the identical task.

Artificial Intelligence··Evening
Two analysts inspect an open computer rig connected by colored cables to a target server rack in a bright, brick-walled test lab.

Claude Mythos alone completes the intrusion chain, at 80 points

The Register reports Booz Allen tested 18 models, nine American and nine Chinese, as autonomous attackers under identical conditions against a defended Active Directory network, scoring what network logs and intrusion-detection sensors proved rather than what the models claimed. Only Claude Mythos completed the intrusion chain end to end on its own, taking administrator control on every attempt with a stolen credential and, on the harder test with no credentials at all, breaking in from outside and still taking the domain; every other model failed that harder test. TNW reports Grok-4.5 followed on 49 and GPT-5.6 Sol on 46, with three models besides Mythos taking full domain control and four more moving laterally inside the network.[1], [2]

A harness turns a 13-point model into an 80-point threat

The Register reports Booz Allen found that Claude Sonnet, listed 15th of 18 at 13 points, rivalled Claude Mythos's 80 once it was paired with an attack harness, the software that connects a model to hacking tools. Booz Allen says it expects most of the other 17 models to reach that same level within six months, and the firm — which also sells cyber defence work — calls mainstream AI attacks imminent, arguing that the US should set resilience deadlines for critical infrastructure.[1], [2]

A second finding: guardrails are not fixed

TNW reports a second, less-noticed finding from the same release: one model declined a task on the grounds that it lacked credentials, while its cyber-tuned sibling, handed the identical task, complied and carried it out. Booz Allen draws a general rule from that, saying guardrails are not a fixed property of a model and their effectiveness shifts with context and configuration, so a refusal in one setting tells you nothing about another. TNW also notes Booz Allen published the index alongside Vellox Labs Guile, a defensive product the firm says cut attacker success by more than 95 percent in its own unverified testing, and that OpenAI's Astra was not tested in the index even though OpenAI said the same week that Astra had reached the company's own critical cybersecurity threshold.[2]

References

  1. News sourceThe RegisterClaude Mythos leads Booz Allen's Cyber Weapon Index with a score of 80↩1↩2
  2. News sourceThe Next WebBooz Allen's Cyber Weapon Index gap closes with a cheap attack harness↩1↩2↩3