Two methods at the same file size
NVIDIA's 17 August post walks step by step through how the 4-bit NVFP4 version of Nemotron 3.5 Lightning was produced. The method has two stages: post-training quantization draws a student out of the full-precision teacher, and the student is then trained against the frozen teacher with a KL divergence loss. The comparison has a real strength: both routes produce the same 21.19 GB file and sit in the same memory against a 65.85 GB full-precision version, so whatever separates them comes from the method.[1]
The number that leads is median score recovery: the middle value among the individual recovery rates on a list of benchmarks. On one intermediate checkpoint, post-training quantization gives 96.33 per cent and distillation carries that to 99.72 per cent, improving 10 of the 11 benchmarks. This is the cleanest table in the method's favour; by the post's own account the quantization here was deliberately pushed harder, leaving a gap for distillation to close.[1]
The individual scores beneath the median
On a second intermediate checkpoint the median rises from 95.84 per cent to 98.53 per cent, yet only 5 of the 9 benchmarks improve. Among those that fall, GDPval Norm. Elo drops from 16.46 to 12.90 and AA-Omni Non-Halluci. from 69.02 to 64.67. The post does not hide this; it says plainly that the median rather than any individual score carries the result. That sentence also describes what the median is here: a summary that folds the gains and the losses on the list into one figure.[1]
The two intermediate checkpoints were not even scored on the same list. The first uses a set of 11 benchmarks, the second what the post calls an updated evaluation suite of 9. On that footing, 99.72 per cent and 98.53 per cent sit on different scales: each reports the middle of its own set. Reading a method's strength off the gap between medians requires treating the composition of the set as fixed.[1]
Where does the number land on the shipped checkpoint?
The real test is the version that reaches users. There the quantization was chosen more conservatively and post-training quantization already sits close to the full-precision baseline: median recovery of 99.24 per cent. The distilled version comes in slightly under that, at 98.97 per cent. On this checkpoint the method's gain does not show up in the median at all; it collects in margins of 3.79 points on Terminal-Bench v2.1, 1.07 on SWE-Bench Multilingual and 0.65 on HLE.[1]
In July this column looked at a benchmark claim in an investor document, where the problem was that the set, the versions and the scoring were never written down. Here all of that is published, and that is real progress. What remains is a question about the measure itself: a single median from a single training run folds improving and declining scores into one number and says nothing about the spread across repeated runs. What would move this method up a rung is per-benchmark results from several runs on the shipped NVFP4 checkpoint and a repetition from outside NVIDIA; the direction of the difference on the hallucination and agentic scores depends on exactly that.[1], [2]