The missing number arrived
Alibaba published the open weights for Qwen3.8-2.4T-A95B: 2.4 trillion total parameters and 95 billion activated per token. The same day, NVIDIA reported that on a GB300 NVL72 rack in FP8 precision, without additional tuning, the model produces over 4,000 tokens per second per GPU and over 350 tokens per second per user. Both rates are vendor measurements and both belong to day one.[1]
On August 3 this column wrote that the Qwen3.8-Max announcement stopped at a total parameter count, and that without an active-parameter breakdown the size of a deployment could not be narrowed. That missing number is now published. The figure that sets how much hardware a deployment needs is the 95 billion parameters that move per token; the disclosed 2.4 trillion total does not give it on its own.[2], [1]
Two rates, one rack
The two published rates describe two different limits. Over 4,000 tokens per second per GPU is a throughput measure, and over 350 tokens per second per user is a latency measure. They are shown on one Pareto curve, yet the post does not state the batch size, the input and output lengths or the reasoning depth at which that point was reached. Without those three numbers, how many concurrent requests the rack was serving cannot be recovered.[1]
The unit being measured here is the rack, not the chip. The GB300 NVL72 gathers 72 Blackwell Ultra GPUs into a single NVLink domain with all-to-all bandwidth of 130 terabytes per second. In a fine-grained mixture of experts the experts chosen for each token are spread across GPUs, so keeping that traffic inside the rack is a precondition for the rate. Even so, it is not settled that the rack network is what sets these numbers: NVIDIA writes in the same post that NVFP4 precision is expected to do better over time, which suggests the day-one ceiling may sit in the software stack as well.[1]
A followable condition comes out of this. If one of the serving paths NVIDIA lists — SGLang, vLLM or NVIDIA Dynamo — publishes the batch size, the input and output lengths and the reasoning setting for this model alongside its tokens per second, the per-user rate becomes a number comparable across stacks. Whether that appears by December 31, 2026 is a signal a reader can check without help.[1]