The route outside GPU memory

SpAx locates a concrete wait inside a language model that exceeds GPU memory: weights repeatedly travel from system memory or storage. The relevant unit is bytes moved for each generated token. The new preprint assigns each weight column to an original, two-bit approximate or unread tier according to activation magnitude. The middle tier preserves contributions that a binary read-or-skip decision would discard.[1]

The physical boundary sits between the GPU and external memory. The researchers used llama.cpp with an RTX 4090, PCIe 4.0 x8, 128 GB of system memory and NVMe storage. Prompts were processed densely; selective transfers applied during decoding. Comparisons use dense decoding of the same model in the same weight format. That isolates the transfer strategy from a change in GPU hardware.[1]

System-memory transfers and storage reads are different operations. SpAx places jointly accessed columns together and merges adjacent read ranges. Coarse storage granularity can move more bytes than a selected column needs. Reconstructing two-bit approximations also consumes GPU time. The format reading the fewest bytes therefore does not always run fastest; the gain depends on transfer time plus reconstruction time.[1]

The denominator of useful speed

My implication for a local-model user is that changing the weight-transfer policy can ease one part of an oversized model's burden. An alternative bottleneck is attention state or other computation under tighter GPU-memory limits. The experiments retain embeddings, the output head and the attention cache on the GPU. Their benefit cannot be assigned unchanged to a setup that must also offload those components.[1]

Quality travels alongside that byte ledger. SpAx evaluates language-model perplexity on WikiText-2 and accuracy on separate downstream tasks. Skipping or approximating weight columns is a lossy intervention. An acceptable quality tolerance changes the denominator of useful generation speed. Fewer transferred bytes alone do not establish more answers at unchanged accuracy.[1]

The concrete engineering milestone is a research implementation operating at batch-one decoding. A deployment account needs the model components retained on the GPU, bytes actually moved per token and time spent reconstructing approximations together. For NVMe, selected bytes and bytes physically read deserve separate entries. SpAx demonstrates a way to feed existing hardware differently, within the measured transfer path and quality boundary.[1]