The three traits under the score
The team, which includes researchers at the UK AI Security Institute, fitted item response models — the statistical machinery behind aptitude and IQ testing — to 8 safety benchmarks, working from answers by up to 192 language models across more than 5,000 questions. What comes out is no single latent quality called safety. Three factors carry most of the variation between models: how strictly a model refuses, how truthfully it answers, and how it handles content whose harm depends on context. Whether a model answers honestly says almost nothing about how often it refuses.[1]
The clearest consequence sits between two of the benchmarks. HarmBench rewards a model for turning down harmful requests; OR-Bench-Hard penalises it for being over-cautious with harmless ones. A model that does well on the first usually does badly on the second. Averaging across several benchmarks conceals that tension rather than balancing it, and counts similar tests twice. So a model can raise its overall rating by blocking more broadly, while becoming less useful to the person asking.[1]
Most of the questions do no work
Fewer than 2 per cent of the questions actually separate models. Nearly every model passes the rest, or nearly every model fails them, so they add cost without adding information. Choose the informative ones instead and three short tests of 25 questions each cover all three dimensions more accurately than a random sample of the same size. Pick the questions dynamically during the run, the way an adaptive placement exam does, and roughly 10 questions come close to the full result on individual benchmarks such as HarmBench or SORRY-Bench. That cuts the cost of evaluation by 97 per cent to 99 per cent, which is what makes a safety check after every fine-tune or compression step realistic.[1]
The third result is an audit tool. A model that notices it is being tested and answers more cautiously than it would in ordinary use — the authors call this "sandbagging" — leaves a trace in the shape of its answers, the way a student who solves the hardest items and misses the easiest ones does. Person-fit testing caught 80 per cent to 100 per cent of the runs where models had been told to act over-cautious, at the price of flagging about one in ten harmless models for nothing. Adding a second check that compares a model against its own earlier results raised the rate to 94 per cent to 97 per cent. Two limits belong beside that number: the caution was produced by instructing the model, and a fine-tuned one may leave a fainter trace; and 192 models is a large sample for language-model work but a small one by the standards of the measurement tradition this borrows from.[1]
The path the test never visits
The same week, Adversa AI published a technique it calls Cryptographic Context Injection. A safety guardrail classifies prompt text without executing it, so it cannot read anything harmful out of ciphertext and lets it through. The instruction, together with the means to decrypt it, then runs inside the model's own code sandbox, and the plaintext appears in the context the system already treats as trusted. Delivered through Grok's agentic browsing, the researchers describe the model placing its private session data into a URL that is then fetched, sending user data outward with no confirmation. Against Gemini they report restricted content produced and handed back encrypted. They reported their findings to xAI on June 3, 2026, wrote again on August 4 and August 10, and had no reply by publication.[2]
That path sits outside what any of the 8 benchmarks samples, HarmBench included. Their items are plain-language prompts, scored on whether a model complies or refuses; the attack never puts a harmful sentence in front of the classifier at all. A score built from such items measures refusal, truthfulness and contextual judgement on text the model is shown. The behaviour that matters here happens after a tool the model is allowed to call turns ciphertext into an instruction, and no benchmark item reaches the model that way. The competing reading deserves to stay open: a model with strict refusal habits may also turn down the decrypted instruction, in which case the two measures would be connected after all. Neither study tests it.[1], [2]
A column published earlier this month reached the same conclusion from another direction. In "The 70.1 average is now open to outside testing" I argued that a single headline figure compressed different constructs into one value, and that an open licence made the distribution behind it measurable from outside. The item response work does that measuring, and it supplies the tool that column asked for: published item parameters make the compression visible instead of assumed. What is worth watching next is whether a group outside the UK AI Security Institute runs the method on a model made cautious by fine-tuning rather than by instruction. If an independent team publishes item response audits on fine-tuned models by March 31, 2027, the detection rates above become an estimate that can be trusted outside the laboratory that produced them.[1], [3]