Eigen RadarAI
Analysis

SpAx reads fewer external weights to speed up local language-model decoding

SpAx divides language-model weights into original, compressed and skipped groups as each token is generated. The single preprint tests how this reduces transfers from system memory or storage when a model does not fit inside GPU memory. Its llama.cpp experiments use one RTX 4090 workstation and distinguish saved transfers from the time needed to reconstruct approximate weights.

Artificial Intelligence··Evening
An open metal computer case seen from above contains a graphics card, memory modules and a motherboard-mounted storage module.

Each generated token changes which weights are read

SpAx addresses a model that must repeatedly fetch weights from outside GPU memory while generating an answer. Its single preprint introduces three reading choices: omit columns associated with the smallest activations, retrieve compressed approximations for intermediate ones, and keep the original format for the largest. An activation is a value produced inside the network. As those values change, the chosen columns can change with each token, without retraining model weights.[1]

A workstation tests memory and storage separately

Memory transfers and direct storage reads are tested separately on a workstation running a llama.cpp-based implementation. Its RTX 4090 is connected over PCIe 4.0 x8; system memory totals 128 GB, alongside NVMe storage. Decoding handles one example at a time. The input prompt uses dense computation, with selective access reserved for output tokens. Calibration sets reading budgets for matrix groups. Combining neighboring selected ranges limits the number of small storage requests.[1]

Fewer bytes and shorter latency are separate measurements

Transferring fewer bytes does not invariably yield the shortest decoding time. Timing includes GPU work to reconstruct approximate weights, and storage request granularity also differs between formats. Reconstruction overhead is therefore a possible explanation for the difference; storage granularity is another, rather than a universal causal ranking of formats. WikiText-2 evaluates language-model probabilities, while downstream tests check accuracy. Embeddings, the output head and the attention cache remain on the GPU. These results cover the tested workloads and quality tolerances.[1]

References

  1. News sourcearXivSpAx reduces offloaded-weight transfers with selective approximation↩1↩2↩3