The rung the rack has reached
The concrete milestone in CoreWeave’s Vera Rubin NVL72 announcement is Cognition running production inference on the system. The cluster was set up in early September, and Cognition engineers tested their SWE-2 workload against a GB200 NVL72 cluster. That moves the claim from a shipment plan to a customer workload. Access remains limited, though: one operating customer cluster does not establish ready, reserved capacity for every cloud customer.[1]
Cognition reported 4.8 times the total token throughput on SWE-2 inference and 3.8 times the output-token throughput on a reinforcement-learning workload. Those are different denominators attached to different jobs. The first compares rack generations on the coding agent’s inference load. The second belongs to another execution pattern. Neither is a universal Vera Rubin speed figure. CoreWeave’s separate DeepSeek R1 result per megawatt is a third experiment and should be read on its own boundary.[1]
From measurement to capacity
Running GB200, GB300, and Vera Rubin clusters under the same management tools may be an operational advantage for CoreWeave. A customer able to use new hardware without rebuilding its toolchain may shorten deployment time. There is another plausible explanation: Cognition’s own workload and preparation may have made this particular transition unusually quick. The release does not measure whether other customers reproduce the same timing or performance.[1]
The next measurable rung is more customers entering sustained service on Vera Rubin with defined workloads and access expanding beyond the limited cluster. An operating rack has now been announced. The release does not say how many customers a rack serves, at what latency target, or for how long. The 4.8 times result therefore matters as Cognition’s reported production test. It cannot stand in for an efficiency ledger covering the cloud provider’s entire fleet.[1]