Shorter decoding, longer action
In BitNest’s OpenVLA-7B experiment, decoding for the same robotic action falls from 135.6 milliseconds to 100.8 milliseconds. Yet total action time rises from 201.5 milliseconds to 212.7 milliseconds. In the researchers’ own measurements, the faster component does not accelerate the complete job. Visual encoding is excluded from both totals, so the gain disappears even within that disclosed boundary. For a control system, the relevant comparison is the combined wait for preparing and generating an action, rather than just the rate of producing its final six tokens.[1]
Prefill rises from 65.9 milliseconds to 111.7 milliseconds. The paper explains the mechanism: small-matrix decoding reads packed low-bit weights directly, whereas the current large-matrix prefill path dequantizes weights to FP16 and uses cuBLAS. Quantization and other transformations remain additional work. A representation that reduces memory traffic therefore does not deliver the same benefit in every computational phase.[1]
Keep the memory gain separate
The latency result does not erase BitNest’s memory gain. Drafting reads a 4-bit base and verification adds a refinement plane to recover an 8-bit representation. No separate draft-weight copy is stored. On Jetson Orin NX 16 GB, LLaMA-2-7B-32K has reported PyTorch peak allocations of 13.9 GiB with FP16, 8.2 GiB with W8A8 and 7.7 GiB with BitNest. System-wide peak memory is reported separately in that experiment. Equating allocated model memory with all device memory would conceal that distinction.[1]
There is more than one useful denominator on Jetson: FP16 is a different precision point, while W8A8 is already a compressed baseline. The researchers report decoding rates of 6.0 and 7.8 tokens per second respectively, with BitNest at 9.0–9.3. For a memory-constrained device, the latter comparison is more directly useful because it separates quantization savings from the additional drafting gain. An alternative explanation is that acceleration comes partly from tuned memory-access kernels rather than nesting alone; the paper separately measures that kernel improvement.[1]
The decision follows the device and workload
For a team choosing hardware, the study provides two separate engineering outcomes: longer executable context within limited memory, and total latency per action. The larger context on Jetson does not remove OpenVLA’s prefill overhead; slower robotic actions do not invalidate the language-model memory savings. BitNest is a preprint and these measurements belong to the researchers’ setup. An implementation decision rests on comparing the same complete workload on the same device. In this example, the bottleneck moves from memory traffic to prefill.[1]