Eigen RadarAI
Analysis

The system around the agent is deciding what it can deliver

Three reports show that agent performance rests on more than model size: a narrow decision boundary, an orchestrator that calls other tools, and a retrievable set of procedures shape the result.

Artificial Intelligence··Midday
A central mechanism routes tasks through bounded apertures to separate tools and a small procedure rack.

Companies finding value are narrowing the decision boundary

According to figures reported by VentureBeat, Gartner expects more than 40 percent of today's agentic AI projects to be stopped before 2028 and counts only about 130 genuinely autonomous products among the thousands marketed as agentic. McKinsey's 2026 survey puts average responsible-AI maturity at 2.3 out of 4. About 30 percent of organisations reach level three or higher in governance and agent controls, while use of agents is scaling roughly eight times faster than governance improves. Nearly two thirds call security and risk the biggest obstacle to expansion. The companies finding value in the report narrow the decisions a tool may take alone.[1]

A smaller model directs the research while a larger tool codes

Faraday, an agent from the London laboratory Inherent, runs on Qwen 3.6, which has 27 billion parameters, and calls OpenAI's GPT-5.5 Codex for coding. The company says Faraday reproduced findings from published papers more faithfully than Claude Opus 4.8 and GPT-5.5 without receiving the answers in advance; the report includes no independent replication. Inherent trained the agent with reinforcement learning that rewards outcomes rather than prescribing every step. The company, whose four cofounders include three former Google DeepMind researchers, emerged from stealth weeks ago with a seed round of 50 million dollars. The model directing the research flow and the tool writing the code are separate components in this arrangement; their division of work is part of the architecture described by the company.[2]

Procedures help, but a crowded library breaks retrieval

Researchers from Princeton, UC San Diego and other institutions ran 8,135 controlled trials and attributed 65.7 percent of the benefit agents obtain from skill files to procedural grounding: setup steps, tool order and intermediate checks. Direct knowledge transfer accounted for 4.5 percent. In the same experiments, agents mechanically applied a playbook that did not fit the task in 10 percent of cases. Access to the right skill also deteriorated sharply as the library grew; retrieval precision fell from 29.6 percent with five skills to 3.3 percent with one hundred. The researchers recommend managing the creation, retrieval and application of skills as one lifecycle instead of accumulating more files.[3]

References

  1. News sourceVentureBeatThe companies getting value from AI agents are narrowing what the agents may decide↩
  2. News sourceTechCrunchThe small model gives the orders and GPT-5.5 Codex writes the code↩
  3. News sourceThe DecoderSkills give agents a procedure, and retrieval collapses once there are 100 of them↩