NVIDIA on 18 September 2026 said AIPerf is not a reskin of Perf Analyzer; the client was rebuilt so Python’s lock does not cap concurrency. It lists more than 15 endpoint types, ShareGPT-style public datasets, and replay traces from Mooncake, Baseten and WEKA AgentX. Core metrics include first-token wait and tokens out per second. The post cites no outside lab bake-off.
Artificial Intelligence··Morning
A new client, not a skin
NVIDIA’s developer blog on 18 September 2026 says AIPerf is not a reskin of Perf Analyzer; the client was rebuilt so Python’s lock does not cap concurrency. That is the vendor’s reason for a new name. TechCrunch, The Verge and Reuters did not file their own story of this rewrite. The card stays on the NVIDIA post. The date is the blog stamp. The bounded search found no second newsroom; the card stays on this announcement.[1]
Endpoints, traces, arrival shapes
The post lists more than 15 endpoint types, public sets such as ShareGPT, and replay traces from Mooncake, Baseten and WEKA AgentX. Steady, Poisson and gamma arrivals with a burst knob let operators shape load, not only volume. Those names are NVIDIA’s feature list. They are not a ranking of which format is best. The bounded search found no second newsroom; the card stays on this announcement.[1]
Metrics, still in-house
Core metrics include first-token wait, the gap between tokens, request wait and tokens out per second, with percentiles and chip telemetry when DCGM or pynvml is installed. The announcement cites no outside lab bake-off. Eigen Radar is not treating AIPerf as a certified benchmark. Those names are NVIDIA’s own dashboard list. The bounded search found no second newsroom; the card stays on this announcement.[1]