Eigen RadarAI
Analysis

As Nvidia touts a perfect public-set run, new measurements show where agents slow down

Nvidia's public-set result highlights an agent architecture at its best, while Amazon and DeepMind show how tool overload, task decomposition and authority design change the outcome.

Artificial Intelligence··Night
In a dark workshop, many different tools crowd the central robot arm's single coupling while a conveyor on the right reaches three illuminated work stations.

Nvidia finished all 183 public levels

Nvidia says its coding agent AVO cleared all 183 levels across the 25 public ARC-AGI-3 environments and reached a 100.00 RHAE score. The company says the result comes from an architecture built around persistent memory, supervision and a repeated inspect-plan-implement-evaluate loop. By Nvidia's own count, AVO completed the same task family with fewer environment actions than the VISTA agent it uses for comparison. But the result covers only the public set, not the semi-private or private competition sets, and the company does not report an outside verification, so for now it remains a vendor measurement rather than an independently settled benchmark result.[1]

The picture changed when six tools became 20

Amazon's SOP-Bench sets up more than 2,000 tasks across 12 business domains to test what happens when an agent has to complete a procedure end to end rather than solve an isolated sub-problem. The sharpest result is that giving the agent more options does not automatically improve the outcome. When the benchmark swapped the six required tools for a 20-tool set, success across 11 frontier models nearly halved. The paper also shows that a newer release can trail an older one inside an agent loop: in the reasoning-style setup, the Claude 4.5 family scored below the older Claude 4 family. Nor was there a universal winner. Accuracy approached 9 out of 10 on easier procedures and fell to roughly 1 in 4 on the harder ones.[2]

DeepMind wants ambiguous work broken apart first

The 'Intelligent AI Delegation' study discussed by Google DeepMind and Google Cloud turns that limit into a design brief. In their account, the orchestrating agent should decompose a goal until each sub-task can be monitored and graded, route work that needs no heavy reasoning to lighter models, grant only the least privilege that works, and preserve the right to challenge or escalate an ambiguous request. That emphasis points at the same operating variables that Amazon measured: how work is split, how many tools are exposed at once, and which requests the agent accepts without challenge. It gives a common language for reading a flashy public-set ceiling and a much messier day-to-day operating limit through the same orchestration questions.[3], [2]

References

  1. News sourceNVIDIANvidia's coding agent clears every level of the ARC-AGI-3 public set↩
  2. News sourceAmazon ScienceHanding an agent extra tools nearly halves its success rate↩1↩2
  3. News sourceGoogle Cloud BlogA DeepMind study wants agents to stop accepting every task without asking↩