Where the Copilot bill fell

In controlled offline evaluations, TerminalBench 2.1 showed a 4.9 percentage point improvement in verified task quality with a 67 percent reduction in estimated cost compared to Claude Opus 5, and CheckpointBench achieved a mean session score virtually tied with the Claude Opus 5 reference baseline with a 65 percent cost reduction. The feature behind those numbers, Project HydraFusion, picks among several models at runtime by task complexity and runs three approaches: a single model for straightforward tasks, a cascade where an efficient model drafts and a quality gate escalates to a stronger model, and a critique where an independent critic model reviews the draft before a structured revision.[1]

To me the interesting question is which of those three produced the saving, and the published figures do not answer it: they compare the assembled system with one baseline model, and GitHub produced them in its own controlled offline evaluations. My reading is that the cascade's quality gate is doing most of the work, because a gate that keeps easy tasks on the efficient model is the only one of the three that removes expensive tokens rather than adding them. That is an inference, not a measurement. The same difference could come from the evaluation set's task mix, where straightforward work an efficient model already solves would lower the estimate with no routing benefit at all.[1]

The Iris release makes the scaffolding checkable

The team behind Iris-mini and Iris-pro tests every benchmark with and without context management while keeping tools, context limits and the judge model constant, and says results reported only with management turned on cannot be cleanly split into what comes from the model and what comes from the scaffolding around it. The size of that share is not small: context management boosts Iris-mini's BrowseComp scores by up to 21.2 points, and the same model reaches 82.2 on BrowseComp with management on while Iris-pro reaches 88.6. The scores also come from a single agent, with no helper agents and no extra verification steps at the end.[2]

What makes that ablation usable by someone else is the second half of the release. The model weights are available in a collection on Hugging Face and the code is on GitHub, and the package includes the Iris Harness with the agent loop, tools, context management strategies, and all four benchmarks with evaluation. The harness runs against any OpenAI-compatible endpoint, so the loop that produced the numbers is the loop a reader can point at a different model. For a team choosing between open-weight agents, that is the part with option value: the weights alone would let you host the model, while the harness lets you reproduce the comparison.[2]

What a builder can re-run

Both releases move the gain into the layer around the model, and they differ in what they hand a reader to check it with. The paired runs that separate the scaffolding's share from the model's appear in the Iris benchmarks, while the Copilot figures describe the assembled system against one baseline. Neither team is hiding anything; they are answering different questions. A product team shipping a research preview into the GitHub Copilot CLI is asked whether the whole thing is cheaper, and a group publishing weights is asked which part of its stack earned the score.[1], [2]

The earlier reading of a coordinated pool of models found that the company shipping the change also produced the measurement, and the Copilot figures repeat that shape. What is new today is that the second release hands over the means to re-run the comparison rather than only the result, which narrows the gap between a vendor number and a number a team owns. A team that runs the Iris Harness on all four benchmarks against its own OpenAI-compatible endpoint, with context management on and off, would turn the scaffolding's share into a figure it produced itself.[1], [2], [3]