What fills the 283 seconds?

NVIDIA deliberately killed one process in a two-worker GLM-5.2 deployment. The model was quantized to NVFP4 on B200 nodes; the synthetic load carried 32,000 input and 1,000 output tokens per request at 0.7 requests a second. The cold path reloads weights into high-bandwidth memory, compiles kernels, and recaptures CUDA graphs. The measurement therefore tests a process fault on healthy hardware under one disclosed workload.[1]

The second worker returned after 283 seconds with a cold restart and after 7.3 seconds with the standby. In the post-fault window, median time to first token fell from 23,815 milliseconds to 1,311 milliseconds, while median decode rate per user rose from 12 to 46 tokens a second. These are NVIDIA's own test results, and the gain belongs to a two-worker setup in which the survivor carried every request during the outage.[1]

What the standby does and does not pay for

The price of the speed is a second inference engine initialized in advance on the same GPU. NVIDIA's GPU Memory Service keeps the physical weight pages independent of the process, so the active and standby engines map the same weights. The standby carries neither a second weight copy nor a key-value cache, but it retains its CUDA context, captured graphs, and communicators. That avoids the two largest memory items while leaving the parked process's supporting state resident on the GPU.[1]

The feature is in preview and requires Kubernetes 1.34 or newer with dynamic resource allocation enabled; vLLM is the primary supported backend. Its boundary is also narrow: it targets process crashes, recoverable CUDA errors, and transient communication failures. Hardware, node, and multi-node failures still rely on standard rescheduling. Valuing the design requires an operator to measure its extra GPU memory and the change in useful throughput across different traffic mixes.[1]