Eigen RadarAI
Analysis

TaSQ compresses model caches while trading initial delay for throughput

The TaSQ preprint reshapes quantization to compress language-model key–value caches to roughly one bit per channel. In the authors’ RTX 6000 Ada serving experiment, larger batches accompanied a reported 1.87-fold increase in peak throughput. Encoding increased time to the first token by 10–14 percent, separating sustained-generation gains from the initial delay. Results remain specific to the tested models and setup.

Artificial Intelligence··Night
A single professional blower GPU connected inside an open aluminum test fixture in a bright laboratory.

Higher throughput comes with an initial delay

The TaSQ preprint reports a 1.87-fold increase in peak generation throughput in the authors’ serving experiment using SGLang on a single RTX 6000 Ada, with the comparison running on the same stack. Peak throughput reached 412.6 tokens per second against 220.8 in the comparison, while maximum batch size increased from six to 84. Encoding the quantized cache increased time to the first token by 10–14 percent on prompts of 8K–32K tokens.[1]

The method reshapes cache quantization

TaSQ compresses the key–value cache that retains earlier-token activations for repeated access during generation. At roughly one bit per channel, a small codebook must represent large channel groups. Query-guided weights, normalization across attention heads and covariance-aware grouping reshape that representation. The authors say the transformations can be incorporated into codebooks and projection weights, preserving the usual lookup structures for vector quantization and compatibility with rotary position embeddings.[1]

Tests separate different model tasks

The single preprint evaluates coding and mathematics separately from knowledge and long-context retrieval. General tests include Qwen3-4B alongside Llama-3.1-8B-Instruct, with additional reasoning models in reasoning experiments. RULER retrieval tests span 4K–64K contexts, using 200 examples and three seeds per setting. The main implementation keeps standard contiguous grouping for values because extra transformations offered limited gains. The reported efficiency and quality results remain the authors’ experiments on the tested models and deployment setup.[1]

References

  1. News sourcearXivTaSQ reshapes quantization space for a one-bit cache↩1↩2↩3