One endpoint, a pool behind it

Sakana AI shipped Fugu Max and Fugu Ultra v2 through the same OpenAI-compatible endpoint it already serves, and says a running deployment moves across with a single-line parameter change. The interesting part is what sits behind that endpoint: a swappable pool of open and specialized models, NVIDIA Nemotron among them, with the router picking the leanest member that can finish the task. Fugu Max is priced at 2 dollars per million input tokens and 6 dollars per million output tokens.[1]

The scores come with a caveat the reader has to supply. Sakana AI reports that Fugu Max takes the best overall score on six benchmarks, among them Terminal Bench 2.1 and the company's internal SWEFish, and that Fugu Ultra v2 reaches 48.3 on Chartography against 27.3 for Opus 5 and 29.5 for Fable 5, and 74.3 on DeepSWE. Every one of those numbers was produced by Sakana AI on a task selection Sakana AI chose, and the announcement offers no independent replication.[1]

The same move inside code review

GitHub applied a narrower version of the same architectural choice. Copilot's Lite effort level stopped running one agent and now runs an ensemble whose findings are merged into a single review. The review agent reaches the full set of shell tools from the Copilot SDK behind the Copilot agent firewall, so it can build, run tests and execute targeted scripts against the code it is judging. GitHub reports that the ensemble raised addressed comments per review by 47 percent for high-severity findings, 31 percent for medium and 11 percent for low, and cut review cost by about 8 percent.[2]

Put the two side by side and the credited mechanism is the same: a coordinated group of models takes the place of one model, and the group gets the credit for the gain. The evidence is the same too, and that is the problem. Sakana AI ran its own comparison and GitHub ran its own experiment; neither publishes which member handled which task, how the task mix was held constant between before and after, or what the failures looked like. A developer reading either number cannot tell whether coordination did the work or whether extra compute per task did.[1], [2]

What a developer can check

Five days ago I argued from Portal's worker modes that the bill for a coding agent is dominated by reading and writing files, so the routing layer is where a team controls cost. Fugu Max's price list is consistent with that view, and it also shows what the view costs: when a swappable pool answers wrongly, the team debugging it has to work out which member answered, and the router's choice is not visible to the caller.[1], [3]

There is a test that would settle it. If an independent evaluator holds the task set and the scoring method constant, publishes per-request routing traces for Fugu Max and reruns the Copilot Lite comparison on a fixed repository set by 31 December 2026, the share of the gain owed to coordination can be separated from the share owed to the pool's composition and to extra compute. Until a comparison of that kind exists, the honest reading of both announcements is that a coordinated pool is a plausible mechanism with vendor-measured support.[1], [2]