Eigen RadarAI
Analysis

Compression tests separate losses in mathematics, code and question answering

A single preprint compares pruning, quantization and distillation by the capabilities a language model must preserve. Similar resource savings can leave different losses in mathematics, code and question answering. The researchers test predictive relationships across model settings, while finding that numerical selection adds little over a simple method priority for question answering within their candidate sets.

Artificial Intelligence··Evening
Gloved hands install a graphics card inside an upright open computer case.

Equal savings leave different capability losses

Reducing a language model's memory footprint can affect its tasks unevenly. A single preprint compares removing weights, representing them with fewer bits and training a smaller student, known as pruning, quantization and distillation. It measures mathematics, code generation and multi-hop question answering separately. Each capability uses 64 fixed probes per model, allowing compression settings to be compared against a consistent set of questions.[1]

A compact predictor still needs new coefficients

Pythia checkpoints let the researchers vary training stage separately from model size. OLMo-2 and a second pruning rule provide transfer tests. The predictor uses a common response shape at different pruning densities: for mathematics and code, this cuts the number of configuration measurements in half. Each setting still needs newly fitted coefficients. Evaluation follows reference-completion probability, with task correctness and executable code checked separately.[1]

Selection gains depend on the task

For question answering, a fixed ordering of compression methods matches numerical selection on the study's regret measure within the tested candidates. Opportunities for extra gains are smaller in mathematics and code. Distillation experiments also find a recurring cost from heavy reuse of training data, while the net benefit changes with the evaluation distribution. The study therefore distinguishes predicting a loss from demonstrating that the prediction improves a deployment choice.[1]

References

  1. News sourcearXivLLM compression study tests predictors of capability loss↩1↩2↩3