Two dimensions without labels

The Allen Institute for AI released BenchMIRT, a multidimensional item response theory method that reads a benchmark one question at a time. Training used results from 100 language models across 16 benchmarks and more than 34,000 questions. Six of those benchmarks measure general reasoning, among them MMLU-Pro, GPQA, MATH and BBH. Ten measure safety, including HarmBench, StrongReject, WildJailbreak, BBQ, WMDP and XSTest. For each model the method estimates strength on the capabilities in that set; for each question it estimates how difficult the item is and how well it distinguishes stronger from weaker models on those capabilities. The technical report, datasets and code were published with the method.[1]

The method was not told which benchmarks measured which abilities. BenchMIRT recovered two dominant dimensions on its own: safety and general reasoning. When the analysis was repeated from scratch, the same two dimensions emerged each time, the authors write. Every model in that training set was released in March 2025 or earlier. The authors also write that a different mix of evaluations could surface different capabilities. Those two limits belong to the design: the dimensions describe this set of 16 benchmarks and this model cohort.[1]

Where the name and the items part

On many suites the method confirmed the intended focus: reasoning benchmarks tracked reasoning, and jailbreak and harmful-content benchmarks tracked safety. Three cases labelled as safety split the name from the items. BBQ, built to test social stereotypes, aligned much more strongly with general reasoning. WMDP, which tests dangerous dual-use knowledge in biology, chemistry and cybersecurity, associated more strongly with general reasoning than with safety; stronger general reasoning sat with lower WMDP scores, and the suite counts refusing or failing to provide that knowledge as the desired response. Inside HarmBench, standard and contextual harmful-request questions aligned more closely with safety, while copyright questions aligned more closely with general reasoning.[1]

Keeping only 10 per cent of a benchmark's questions generally preserved the capability ranking; keeping 50 per cent often tracked the full benchmark more closely. On held-out questions the method predicted performance with 79 per cent accuracy against a 70 per cent baseline that assumes a model does about as well on each question as on the benchmark overall. The authors also note a trade-off: if the goal is to rank models by predicted performance on randomly held-out items, the benchmark's average score performs slightly better than BenchMIRT; the method's advantage is the finer picture of individual questions. A ranking that survives on 10 per cent of the items still ranks whatever mixture those items contain. BBQ remains a reasoning-aligned suite in that ranking, and WMDP's dual-use knowledge stays unseparated from the reasoning used to refuse it.[1]

A second split inside a knowledge claim

Researchers at Google Research and Technion tested 13 language models on WikiProfile, a set of 2,150 Wikipedia facts asked in several formats across more than 4 million responses. Frontier models such as GPT-5 and Gemini-3 encoded 95–98 per cent of the tested facts and did not directly state 26–34 per cent of what they had encoded. Giving the models more thinking time at inference recovered 40–65 per cent of the encoded facts they had missed, while only 10–20 per cent of facts needed that extra thinking at all. The gap is uneven: in frontier models rare facts sat more than 25 per cent below popular ones, and reversed questions were harder than the direct form. The measurements are the authors' own; no replication has been published.[2]

BenchMIRT and WikiProfile show the same kind of split under a different name. A safety score and a knowledge claim both arrive as one construct; the items, or the prompt form, carry at least two. BenchMIRT's two dimensions and WikiProfile's encoded-but-unstated facts share that common constraint. An alternative, plainer reading remains open: the two dimensions may simply mirror the mix of six reasoning and ten safety benchmarks that was fed in. The WikiProfile gap may be format sensitivity rather than a store of facts the model holds; reversed questions being harder is consistent with that. The next useful measurement is specific. For BenchMIRT, recover the dimensions on a mix that is not six reasoning suites plus ten safety suites, and on models released after March 2025. For WikiProfile, an independent run of the same 2,150 facts would show whether the 26–34 per cent unstated share appears outside the authors' own measurement.[1], [2]