NVIDIA keeps a standby inference engine ready on the same GPU
In NVIDIA's test, a standby inference engine on the same GPU resumed service 7.3 seconds after a fault, versus 283 seconds for a cold restart. A McKinsey survey found agent scaling rose while earnings impact stayed flat. A theoretical study warns that cheaper-model routing under congestion can trigger retries that add demand.
Artificial Intelligence··Midday
The standby engine avoids reloading the weights
NVIDIA's preview Shadow Engine Recovery feature for Dynamo keeps a fully initialized standby beside the active inference engine on the same GPU. The engines share model weights through GPU Memory Service, so the standby does not hold a separate copy. In NVIDIA's synthetic-load test using GLM-5.2 on B200 nodes, a second worker resumed service 7.3 seconds after a fault, compared with 283 seconds for a cold restart. These are the vendor's own results. The feature requires Kubernetes 1.34 or newer with dynamic resource allocation and principally targets vLLM.[1]
Agent scale rises while earnings impact stays flat
The McKinsey survey reported by The Register covers 1,719 professionals and business leaders worldwide. The share attributing at least some earnings impact to AI stayed at 37 per cent, while 6 per cent qualified as high performers; both were unchanged from 2025. Among companies with revenue above 1 billion dollars, the share scaling agents rose from 27 per cent a year earlier to 40 per cent. While 80 per cent of AI users reported higher individual productivity, 20 per cent said running cost limited use. The survey does not measure how any single infrastructure safeguard affects earnings.[2]
Cheaper routing can add demand through retries
The theoretical study models a system in which routing queries to smaller models during congestion produces weaker answers that can prompt users to retry. Once retries count as part of incoming demand, a temporary surge can become a persistent low-quality state. This is not an observed result from a live service; it is a theoretical construction calibrated to public benchmark data. In a real service, retry behavior, queue management or capacity allocation could produce a different result. A standby engine targets recovery time after failure, while this study isolates how model choice under congestion could change demand.[3], [1]