Long context narrows at the attention layer
NVIDIA did not treat long-context latency as a faster-hardware problem alone. The published analysis measures how the ratio of query heads to key-value heads, the head dimension and the sequence length govern inference time on GPUs. In a post signed by Nidhi Bhatia, Samkit Jain, Sai Kishan Pampana, Ritika Borkar and Bita Darvish Rouhani, decode performance scales linearly with group size, giving roughly a twofold speedup each time the group size doubles, while prefill is insensitive to the same parameter. Head dimensions of 128 or 256 are recommended so that the layer aligns with GPU tile sizes. Prefill grows quadratically with input length; decode grows linearly with the length of the key-value cache. The measurements use FP8 precision for both attention compute and the key-value cache, and NVIDIA Nemotron 3, which uses grouped-query attention with two key-value heads, is given as an example. The post gathers tensor parallelism, wide expert parallelism and Helix parallelism as implemented in TensorRT-LLM into four design recommendations. No new model weights or code package were released.[1]
