Eigen RadarAI
Analysis

Search support lifts the score from 33 to 75, while memory does not pay off the same in every model

Newly published hardware and benchmark tests reveal model speed on large datasets and their reliance on auxiliary tools. Search capability and agent memory directly influence result quality at inference, though not every model benefits equally.

Artificial Intelligence··Morning
A vast 3D network sculpture with glowing node clusters hangs in a white hall; a lone researcher beneath.

Eight H100s built a 106 million vector graph in 8 minutes

Testing hardware boundaries, NVIDIA updated its cuML and cuVS library versions to distribute UMAP graph construction across multiple graphics processing units rather than confining it to just one. According to internal company performance measurements, eight H100 data center accelerators process a massive 870 GB dataset comprising 106 million vectors in a mere 8 minutes, achieving high acceleration and efficiency compared to system-projected runtimes.[1]

Without search the score stops at 33

The impact of tool use on model performance is measured by the Search Index published by the Artificial Analysis firm. The conducted tests rank the effect of various search services used by agents during tasks on quality, cost, and speed; a model given no search tool stops at only 33 points on difficult tasks, while providing search capability lets the same model reach up to 75 points and elevate its performance significantly.[2]

Memory does not work in every architecture

The performance gain brought by agent memory also fails to work the same way across every architecture and weight class. IBM Research units tested ALTK-Evolve, a method that feeds behavioral guidelines extracted from an agent's past trajectories back into the model at inference time, on AppWorld tasks; while the gpt-oss-120b model showed a 16.1-point increase in completing complex tasks, the GLM-5 model was found to gain no additional benefit from this memory addition and its performance remained unchanged.[3]

References

  1. News sourceNVIDIA Technical BlogEight H100s build a 106 million vector UMAP graph in 8 minutes↩
  2. News sourceThe DecoderWithout search the score stops at 33; with search it reaches 75↩
  3. News sourceHugging FaceAgent memory does not pay off the same way in every model↩