As Nvidia touts a perfect public-set run, new measurements show where agents slow down
Nvidia's public-set result highlights an agent architecture at its best, while Amazon and DeepMind show how tool overload, task decomposition and authority design change the outcome.
Artificial Intelligence··Night
Nvidia finished all 183 public levels
Nvidia says its coding agent AVO cleared all 183 levels across the 25 public ARC-AGI-3 environments and reached a 100.00 RHAE score. The company says the result comes from an architecture built around persistent memory, supervision and a repeated inspect-plan-implement-evaluate loop. By Nvidia's own count, AVO completed the same task family with fewer environment actions than the VISTA agent it uses for comparison. But the result covers only the public set, not the semi-private or private competition sets, and the company does not report an outside verification, so for now it remains a vendor measurement rather than an independently settled benchmark result.[1]
The picture changed when six tools became 20
Amazon's SOP-Bench sets up more than 2,000 tasks across 12 business domains to test what happens when an agent has to complete a procedure end to end rather than solve an isolated sub-problem. The sharpest result is that giving the agent more options does not automatically improve the outcome. When the benchmark swapped the six required tools for a 20-tool set, success across 11 frontier models nearly halved. The paper also shows that a newer release can trail an older one inside an agent loop: in the reasoning-style setup, the Claude 4.5 family scored below the older Claude 4 family. Nor was there a universal winner. Accuracy approached 9 out of 10 on easier procedures and fell to roughly 1 in 4 on the harder ones.[2]
DeepMind wants ambiguous work broken apart first
The 'Intelligent AI Delegation' study discussed by Google DeepMind and Google Cloud turns that limit into a design brief. In their account, the orchestrating agent should decompose a goal until each sub-task can be monitored and graded, route work that needs no heavy reasoning to lighter models, grant only the least privilege that works, and preserve the right to challenge or escalate an ambiguous request. That emphasis points at the same operating variables that Amazon measured: how work is split, how many tools are exposed at once, and which requests the agent accepts without challenge. It gives a common language for reading a flashy public-set ceiling and a much messier day-to-day operating limit through the same orchestration questions.[3], [2]