Two announcements, two measurement setups
One number carries Microsoft's announcement: MDASH, the harness that hunts for vulnerabilities, reaches 96% on the CyberGym benchmark with the new MAI-Cyber-1-Flash model, 12 points above Anthropic's Mythos, at roughly half the cost of the current configuration. A separate Microsoft AI post opens up what was measured: MDASH running MAI-Cyber-1-Flash together with GPT-5.4, compared against a current configuration of GPT-5.4, GPT-5.4 mini and GPT-5.3 codex. Project Perception, the agent system announced alongside it, enters public preview on August 3. Every figure is the company's own measurement.[1]
NVIDIA published NOOA the same day: an open-source research harness that expresses an agent as a single Python class, with code on GitHub and a technical report on arXiv. Its reported results come from running the same harness with different models: 82.2% on SWE-bench Verified with GPT-5.5 and 79.8% with Opus 4.6; 86.8% on CyberGym L1 with GPT-5.5; 50.2% on ARC-AGI-3 with GPT-5.5 and 85.1% with GPT-5.6-sol. NVIDIA also reports parity or better at roughly half the token cost of the harnesses it compares against, and those figures are equally its own.[2]
What does the comparison hold constant?
Both Microsoft configurations are model mixes. The old mix is three OpenAI models; the new one is a Microsoft model plus GPT-5.4. What was measured is therefore the result after the models inside the harness were swapped, not what the harness itself contributes. Isolating the harness would mean holding it fixed while the models change, or holding the models fixed while the harness changes; here both move at once. The other reading is available too: Microsoft may have done that separation internally and simply left it out of a product post.[1]
NOOA's numbers make exactly that separation, and the result is instructive. On ARC-AGI-3 the harness stays fixed while the model changes, and the score moves from 50.2% to 85.1%. On SWE-bench Verified the gap between the two models is 2.4 points. The same harness makes model choice nearly immaterial on one task set and leaves most of the outcome to the model on another. For a builder the thing to read is that spread rather than an average: on which job does the harness make the model swappable, and on which does it not?[2]
What stays in the builder's hands
In my column of 25 July I argued that with Opus 5 the decision was moving from picking a model to setting effort per request. These two announcements move it again: now the thing being picked is the harness. But the two do not leave the builder the same room. NOOA's code can be downloaded, and its published numbers show what happens when the model underneath is changed; Project Perception opens in preview on August 3, and there is as yet no figure for how MDASH performs with models from outside one supplier. Closedness need not be the only explanation: in a security product, fixing the model can be an engineering decision meant to keep supported behaviour predictable.[1], [2], [3]
One thing would be enough to make it testable. If Microsoft publishes a CyberGym result for MDASH in the same setup with a non-Microsoft model by 31 October 2026, the 96% and the halved cost stop being a claim about one configuration and become a claim about the harness. If it does not, the number is not thereby wrong; it simply stays a number that does not say which layer produced the gain.[1]