Where in the stack the change lands

DLR-Lock replaces every pretrained multilayer perceptron in a model with a deep low-rank residual network of comparable parameter count. Those networks are trained by module-wise distillation and force activation memory that grows linearly with depth during backpropagation. The paper says the resulting architectural mismatch complicates the optimisation landscape of standard fine-tuning, and that the backward pass incurs disproportionately more overhead than the forward pass.[1]

The authors report that the defence preserves the original model's capabilities. The change therefore sits at the model layer and the side that runs the model stays largely where it was; the whole burden falls on whoever differentiates. That distinction matters to a builder, because running a model and adapting it to your own data do not come out of the same budget line.[1]

Who carries the overhead

Activation memory that grows linearly with depth is the first wall a small team hits when it fine-tunes on a single accelerator. A lab with a cluster can spread the same overhead across more memory. What I take from this is that the method ties the possibility of adaptation to a hardware budget more than it makes adaptation impossible. Another explanation is available: if the distilled replacement network also costs something on the forward pass, the burden does not sit only with the adapter.[1]

The limit of that reading is the absence of numbers. The publication page gives no quantitative results, no named benchmark dataset and no overhead multiplier; it says only that experiments on language models validate the claims. How much a defence actually protects is hidden in the size of that multiplier. A method that doubles the backward pass and a method that multiplies it tenfold are two entirely different decisions for the same builder.[1]

The decision this leaves a team with

Until now a team tying itself to an open-weight model asked two questions: can the weights be downloaded, and what does the licence say. DLR-Lock adds a third: what does it cost in hardware to adapt those weights to your own data? Until that answer is published, the set of options an openly released model leaves a builder cannot be measured. I think the centre of the argument moves from licence text to a measurable cost, and it hangs there until the measurement appears.[1]