Compression tests separate losses in mathematics, code and question answering
A single preprint compares pruning, quantization and distillation by the capabilities a language model must preserve. Similar resource savings can leave different losses in mathematics, code and question answering. The researchers test predictive relationships across model settings, while finding that numerical selection adds little over a simple method priority for question answering within their candidate sets.
Artificial Intelligence··Evening
Equal savings leave different capability losses
Reducing a language model's memory footprint can affect its tasks unevenly. A single preprint compares removing weights, representing them with fewer bits and training a smaller student, known as pruning, quantization and distillation. It measures mathematics, code generation and multi-hop question answering separately. Each capability uses 64 fixed probes per model, allowing compression settings to be compared against a consistent set of questions.[1]
A compact predictor still needs new coefficients
Pythia checkpoints let the researchers vary training stage separately from model size. OLMo-2 and a second pruning rule provide transfer tests. The predictor uses a common response shape at different pruning densities: for mathematics and code, this cuts the number of configuration measurements in half. Each setting still needs newly fitted coefficients. Evaluation follows reference-completion probability, with task correctness and executable code checked separately.[1]
Selection gains depend on the task
For question answering, a fixed ordering of compression methods matches numerical selection on the study's regret measure within the tested candidates. Opportunities for extra gains are smaller in mathematics and code. Distillation experiments also find a recurring cost from heavy reuse of training data, while the net benefit changes with the evaluation distribution. The study therefore distinguishes predicting a loss from demonstrating that the prediction improves a deployment choice.[1]