NVIDIA Vera Rubin NVL72 entered the MLPerf Inference v6.1 preview
MLCommons listed NVIDIA Vera Rubin NVL72 in preview in Inference v6.1. NVIDIA wrote preview throughput up to 3.7 times GB300 NVL72 on Qwen3-VL and up to 2.5 times on DeepSeek-R1. StorageReview called those Rubin’s first peer-reviewed numbers.
Artificial Intelligence··Night
Preview list
MLCommons published MLPerf Inference v6.1 results and wrote that NVIDIA Vera Rubin NVL72 is in the preview category. The release adds new tests. Submitter count is higher than the previous round. The NVIDIA submission is dated 16 September 2026. The Qwen3-VL line covers offline, server and interactive scenarios. The DeepSeek-R1 line names TensorRT-LLM. MLCommons kept the preview category distinct.[1], [2]
NVIDIA numbers
NVIDIA published Vera Rubin NVL72’s first MLPerf Inference v6.1 preview submission on 16 September 2026. It wrote up to 3.7 times higher throughput than GB300 NVL72 on Qwen3-VL across offline, server and interactive scenarios. On DeepSeek-R1, with TensorRT-LLM, it wrote up to 2.5 times. The same post named vLLM and NVIDIA Dynamo on the Qwen3-VL line.[2], [3]
Peer review
StorageReview wrote that the v6.1 round carries the first peer-reviewed numbers for NVIDIA’s Vera Rubin NVL72. It said the best per-accelerator DeepSeek-R1 figure for the server scenario is 5.7 times the v5.1 mark from a year earlier. Rubin appears in preview, submitted by NVIDIA and by Nebius. StorageReview also listed AMD Instinct MI350P and Intel Arc Pro B70 numbers in the same round.[1], [3]