The denominator on an A100
An A100 with 40 GB of memory is this study’s physical boundary. In the TopK-Guided preprint, Mukund Agarwalla and Chih-Jen Lin allocate inference computation for Llama-2-7B and Llama-3-8B on that GPU. Their targets are 30%, 50% and 70% activation sparsity. Those percentages describe activations excluded from matrix computation; they do not establish corresponding reductions in the device’s electricity draw or answer latency. The denominator belongs to the operation budget.[1]
Sparsity here does not permanently erase model weights. For each token, some columns associated with near-zero activations are left out of computation. With hardware held fixed, the software intervention is explicit: choosing which computations the same model performs. A result in watts also requires operating-time and power measurements. The preprint’s measurement set is confined to floating-point operation counts and output quality.[1]
Moving the computational cut
Keeping the budget has two separate problems. TEAL’s threshold approach can assign different amounts of computation to different tokens, while WINA’s fixed-count selection preserves the budget with less adaptation. TopK-Guided first estimates the share of activations below the threshold and clips it within one percentage point of the target. Quickselect chooses which activations to retain. Token-dependent selection and a predefined computation limit therefore operate within the same arrangement.[1]
Layers also receive unequal reductions. Normalized error between dense and sparse outputs identifies fragile blocks. More activations remain in sensitive early blocks, while robust later blocks take larger cuts; the total budget stays fixed. The engineering contribution is to preserve computation where its removal causes more error, instead of spreading the same cut everywhere. The gain concerns allocation of existing computation before it concerns capacity from another chip.[1]
From a kernel to service capacity
In the authors’ Llama-3-8B experiment at 70% sparsity, average accuracy across eight tasks is 57.85% for TopK-Guided, 55.34% for WINA and 52.64% for TEAL. The comparison shows allocation affecting quality under the same model and sparsity target. It does not establish the same result for larger models or different workloads. The 7B and 8B models, the GPU and these tasks define the measurement boundary.[1]
The missing link between operation counts and service capacity is an execution kernel. The authors leave a fast kernel implementation for subsequent work and provide no end-to-end latency or energy-savings measurement. Their result therefore cannot calculate how many additional requests a data center can serve. What TopK-Guided demonstrates today is more selective allocation of a fixed operation budget. Separate measurement of the conversion into GPU time determines the contribution to operating expenditure.[1]