Two service objectives on one card

Running four conversations together in Strata redistributes GPU memory. The local inference engine’s developer reports that the final request started after 1.8 seconds rather than 11.2 seconds in a Q2_0 experiment on an RTX 5070. In that same experiment, combined generation speed fell from 70.7 to 63.1 tokens per second. The waiting user received service sooner while the card produced less overall. Those outcomes describe different service objectives on the same hardware.[1]

Each open session takes 0.56 GiB from the expert cache, the GPU space holding expert components used by the model. Shrinking that space to accommodate concurrency imposes a cost on the work already running. The developer reports that even a lone request slowed by 11 percent with two slots enabled and 22 percent with four. The card is unchanged; opening more sessions makes session memory compete with memory allocated to model components.[1]

My operational reading is that a card’s usefulness should be judged using both total token production and the next user’s wait. Earlier starts may matter during a short burst of requests; reduced aggregate output may weigh more heavily under sustained load. The memory explanation gives that choice a concrete mechanism. Software overhead from managing sessions together could also contribute to the slowdown, however; the published experiment does not measure those effects separately.[1]

Allocating memory and delivering hardware support

Strata therefore keeps concurrency optional and recommends it where most experts fit in GPU memory. A long prompt can yield to a shorter request at a chunk boundary, new conversations wait when no slot is free, and a lone request returns to the sequential path. The practical choice is which memory allocation and waiting-time objective to prioritise when opening more conversations on the computer.[1]

Hardware expansion is also at several delivery stages. A CUDA 12.9 engine has been compiled for Pascal and Volta, but the developer has neither card family. The Intel Arc path builds from source on Linux, its kernel tests used a CPU device, and no packaged Intel engine is supplied. Users of these options take an experimental route to an application running on their own card. Completing a build precedes measuring sustained service capacity on that device.[1]

A useful next comparison would hold the card, model and prompts constant while varying concurrent sessions. Reporting combined generation speed, the last request’s start delay and memory allocated to the cache together would show the output cost of a shorter queue. In this Strata release, the central decision concerns allocation of existing GPU memory. A setup prioritising the queue and one prioritising aggregate production choose different services from the same card.[1]