The small model does the directing
Inherent, a London lab founded by DeepMind alumni, says its newly released agent Faraday reproduced the findings of published scientific papers more faithfully than Claude Opus 4.8 and GPT-5.5, without being told the answers in advance. Faraday runs on Qwen 3.6, a model of 27 billion parameters, which leaves it small beside either system it was measured against. What it reaches for when it needs to write code decides how that result was produced: instead of building a coding tool of its own, Inherent had Faraday call OpenAI's GPT-5.5 Codex, the way a researcher leans on software that already exists.[1]
So one of the two systems Faraday was measured against is also the system it uses to do the work. If that comparison holds up, the gain sits in the layer that decides which experiment to run and how to check it, and the model that writes the code becomes a component that can be swapped. I read it that way, and the honest counter-reading is close behind: Inherent trained Faraday with reinforcement learning that rewards outcomes on replication work, so the advantage may come from fitting that task distribution more than from coordination. A director trained on this suite would then carry none of it to work that looks different.[1]
Nvidia's harness raised the same question a day earlier
A day before Faraday, Nvidia published research on Agentic Variation Operators, a harness of memory, tools and a supervisor component wrapped around a model it did not train. With Claude Opus 5 inside, the setup scored 100 percent on ARC-AGI-3, against 30 percent for the model on its own. Adel El Hallak, a vice president in Nvidia's AI product unit, described an agent as the scaffolding around the model together with the set of tools it uses, and argued that open harnesses let teams turn many more knobs to drive up accuracy.[2]
Put side by side, the two results differ in which piece was trained. Nvidia's harness is configuration a team can inspect and adjust — memory, tools, a supervisor — around weights it rents. Inherent's director is itself a trained model, with the large system sitting below it as a tool. That is a sharper version of the same claim, and it arrives with the same limit: in both cases the party that built the layer chose the task suite and ran the scoring. The Nvidia report at least carries one outside number, since TechCrunch notes that OpenAI tripled its own ARC-AGI-3 result by changing two harness settings without reaching Nvidia's score. For Faraday no comparable outside figure is reported.[1], [2]
What a builder actually takes on
In the column on Cloudflare's code review I argued that the leverage arrives before the model, as engineering knowledge turns into rules with named owners. Faraday extends that argument in an uncomfortable direction. A rule set in a review pipeline stays readable and editable by the team that wrote it. A trained director gives the team neither of those. A builder who adopts Faraday can still swap the coding tool underneath it, since Inherent chose GPT-5.5 Codex and could have chosen otherwise. The piece that decides what to attempt, though, sits inside weights the builder did not train, and the report says nothing about whether those weights can be obtained or inspected.[3], [1]
That gives a test worth running in place of a verdict. If the advantage really sits in the director, then holding Inherent's task suite and GPT-5.5 Codex fixed and replacing Faraday with an ordinary prompt-driven coordination layer should reproduce most of the gap; if it comes from fitting the replication work, the gap should shrink on a suite Inherent did not assemble. My conditional is this: if an independent group publishes such a run before 30 November 2026 and the gap survives on tasks Inherent did not choose, the case for a trained coordination layer stops resting on the company's own account. Until that run is published, what we have is one lab's measurement of its own product.[1]