TensorRT 11.0 spreads one model across multiple GPUs
NVIDIA integrated TensorRT 11.0 multi-device support into Dynamo-Triton 26.07, serving a distributed model behind one endpoint. In the company's Cosmos 3 Nano test, a 189-frame video fell from 157 seconds on 1 GPU to 34 seconds on 8 GPUs; model loading and encoding were excluded. AI Brainer confirms the integration, while concurrent-request throughput and total cost of ownership were not measured.
Artificial Intelligence··Midday
One model endpoint now owns several GPUs
NVIDIA added TensorRT 11.0 multi-device inference to the Dynamo-Triton 26.07 serving platform. A distributed TensorRT plan can now execute across several graphics processing units (GPUs) while one named model remains visible to the application through a single endpoint. Dynamo-Triton creates the execution context, stream, and communication resources for each participating rank. AI Brainer independently reports the same integration and its central change: the client no longer has to coordinate the GPU ranks itself.[1], [2]
Cosmos test cuts one video run from 157 to 34 seconds
NVIDIA demonstrated the integration with Cosmos 3 Nano video generation. Its test generated a 189-frame video in 157 seconds on 1 GPU and 34 seconds on 8 GPUs. The same 8-GPU system was used across the configurations, with one warm-up run followed by 5 measured runs. The distributed plan divided the video sequence across the participating devices, while the surrounding pipeline continued to handle prompts, scheduling, decoding, and final video processing.[1]
The test leaves throughput and ownership cost open
The timing excludes model loading and video encoding, and it measures one request rather than a stream of concurrent requests. NVIDIA says the example did not measure throughput under concurrency or total cost of ownership. AI Brainer confirms the multi-GPU integration but does not add an independent performance test. The published result therefore shows lower latency for NVIDIA’s stated Cosmos configuration; it does not establish how the platform behaves under mixed production traffic or whether using more GPUs lowers the cost of each completed request.[1], [2]