The vendor sets the calendar, the buyer does the work
The benchmark, built on four consecutive Qwen releases and six tasks, tests four ways of moving a fine-tuned specialist onto a new base model. The result is sharp: copying the adapter straight across preserves 0.88–0.99 of the gain over a 46 billion-token continuation, but between two independently trained releases it drops by roughly 50 points on the Banking77 task. What sets the distance between two models carrying the same name is training history; architectural similarity does not close it.[1]
That measurement makes a quiet arrangement visible. The model vendor sets the release calendar; the organisation that specialised on the old release has to decide again at every new one, and usually to redo the work. Vendor cadence is not the only possible explanation: a buyer's own product roadmap, cost pressure or compliance requirements can force the same migration. What the study does settle is that the cost of the migration decision depends on the task and on the training distance between the two releases, and the buyer controls neither variable.[1]
There is a saving in the labels, none in the compute
This is where the study draws its sharpest line. Relabelling with the old specialist's own output reaches a retention of 0.96–1.05, and the authors state plainly that it saves annotation rather than compute: the same refresh costs 39.1 GPU-minutes on Banking77 and 255.0 on Spider. The decision policy tested over 33 upgrade episodes, meanwhile, runs at a mean quality regret of 0.37 points against always retraining, while using 33 percent of the training compute and 27 percent of the gold-label episodes.[1]
The two savings do not carry the same weight. Compute sits line by line on a provider's invoice; a gold label means a person sitting down to mark examples, and in most organisations it is not tracked as its own budget line. Since the adapter itself loses its value across releases, what stays durable in the organisation's hands is the task data and the labels the old specialist can still produce. Until the upgrade conversation moves its centre of gravity there, the most expensive input will remain the least visible one.[1]
The line item no meter sees
On 22 August this column argued that the speed of agentic tools books verification and procedure labour into a worker's after-hours time, while the quantities anyone measures are token spend. The finding here shows the same gap repeating one layer up, at the level of the organisation: the metered item is again the machine side, and the unmetered item is again human work. The difference is that this time the number exists; how much of the labelling burden can be removed is a measured share rather than a guess.[1], [2]
There is a signal worth watching. If an organisation really is deciding per task, the trace will show up in the model migrations it completes by December 31, 2026: the share completed without opening a new gold-label order rises. If that share does not rise, the organisation is either not deciding per task or is looking for its saving in compute alone, and in either case the question of where the annotation work gets booked stays unanswered.[1]