Which workload, which library, on the preview
NVIDIA put Vera Rubin NVL72 into the MLPerf Inference v6.1 Closed Division preview. On Qwen3-VL it wrote up to 3.7 times GB300 NVL72 throughput with vLLM and Dynamo; on DeepSeek-R1, up to 2.5 times with TensorRT-LLM. The entries are 6.1-0106 and 6.1-0074. A 288-GPU GB300 run posted 99 per cent scaling efficiency. The post-submission GPT-OSS-120B line is not yet verified. On 13 August two NVIDIA rates lacked a shared denominator; today the denominator exists: named workload, named library, preview rung.[1], [4]
The preview submission reports Closed Division throughput on the vendor's chosen software stack. Seeing the same megawatt on an independent customer rack, at the same latency target, is another rung. Put the figures on one denominator: 3.7 times is tied to Qwen3-VL, 2.5 times to DeepSeek-R1; melting both into one "Vera is faster" line loses the denominator again.[1]
A fabric sketch and a lease do not sit beside a measured rack
In the same 24 hours Cornelis described Active Compute Fabric and Delos Data an I/O die it put at more than 30 Tbps, about 8 TB/s either way. The report parks Nvidia and AMD accelerators at 3.6 TB/s and ties any Delos shipment to integration. The Register frames two sketches entering the scale-up race. Scale-up fabric is the next bottleneck candidate on the path from a measured rack to a wider row of accelerators; shipped silicon is absent from these pages.[2], [1]
Anthropic signed to use part of a 32 billion dollar proposed data centre on Queensland's Western Downs, its first Australia deal, pending FIRB approval. The ABC copy has no megawatts. The lease sits on the contract rung. The next readable event is an independent customer MLPerf reprint on the same Qwen3-VL and DeepSeek-R1 bounds.[3], [1]