Which original?
Multiverse Computing presents Quantization-Aware Healing as a recovery recipe that departs from the usual approach at a single point: it distils the compressed student directly from the original, pre-compression model, held frozen as a teacher. It does this with a KL-divergence loss on logits and a chunked computation that can carry sequences of 32,000 tokens. Applied to GPT-OSS 120B, cut to 60B parameters and quantised to MXFP4, the resulting model beats its own bfloat16 version on 7 of 9 benchmarks. That bfloat16 version is itself a product of the compression: the 60B checkpoint before quantisation, rather than the 120B network the recipe distils from.[1]
The gains in that comparison are real and unevenly spread. Against the 60B bfloat16 baseline the 4-bit model adds 7.4 points on AA-LCR, 5.6 on AIME 2025 and 2.7 on Aider, and gives back 1.4 on SciCode and 0.2 on MMLU-Pro. What the headline does not describe is the comparison a buyer actually faces. That choice lies between the compressed model and the 120B teacher it was distilled from, and the table separates in that column: 66.5 against 66.0 on LiveCodeBench, 67.4 against 69.0 on GPQA Diamond, 42.7 against 50.0 on AA-LCR.[1]
The benchmark that moves most
AA-LCR moves furthest in both directions, and that is where I would look first. It is the long-context reasoning task in this set, the score that depends most on how much of a long input the model can still hold together after its weights are cut and quantised. The recipe recovers 7.4 points of it against the compressed sibling and leaves it 7.3 short of the teacher. The plainest reading is that healing restores much of what quantisation costs on long inputs while the halved parameter count keeps a real gap open. Another reading is available: AA-LCR may be the noisiest item in the set, in which case a repeat with different seeds would move it more than the others. The post reports results from single runs.[1]
This is the second time this month that a quantisation recovery claim has arrived with its comparison set chosen by the vendor. When I read NVIDIA's median recovery rate a week ago, the difficulty was that the composition of the scoring set changed between experiments, so the median moved for reasons other than the method. Multiverse Computing publishes the per-benchmark table NVIDIA withheld, and that is the more useful disclosure. What stays missing in both is the same: the evaluation harness, the number of runs, and any measurement taken by someone outside the vendor.[1], [3]
The measurement that would settle it
The measurement that would settle this is narrow and cheap: the same nine benchmarks, run on the released 4-bit checkpoint and on GPT-OSS 120B under one declared harness, at least three times with different seeds, with the spread reported alongside the mean. Multiverse Computing points to an arXiv paper numbered 2608.20953 for the method; the post itself does not say whether the weights are open, and without them nobody outside the company can run that check. If the weights and a harness are published by 30 November 2026, I would expect the AA-LCR gap against the teacher to hold above 4 points and the LiveCodeBench win to fall inside run-to-run noise.[1]
Comparator choice is doing similar work elsewhere this week. Nvidia said the Groq 3 LPX reached 3,400 tokens per second on Gemma 4 31B where Cerebras sits at 882 tokens per second, though that same figure needs at least 64 accelerators on one side and one or two units on the other, and it leaves out Cerebras's current CS-4 generation. The two cases differ in what a reader can reconstruct. Multiverse Computing prints the teacher row, so the comparison the headline skips can be recovered from the same table; the LPX claim omits the competing generation entirely, and no arithmetic on the published figures brings it back.[1], [2]