New studies measure repeated work in model upgrades and retrieval
Two preprints examine preserving fine-tuning across model versions and the cost of rereading text in retrieval systems. Both focus on reusing work already performed.
Artificial Intelligence··Midday
Adapters do not survive every version change
The preprint on four consecutive Qwen releases and six tasks reports that moving a fine-tuned specialist depends on the transition. According to the arXiv study, copying an adapter directly retains 0.88 to 0.99 of gain over a 46 billion token continuation. Between independently trained releases, Banking77 loses roughly 50 points. The authors also evaluate 33 upgrade episodes. The work has not been peer reviewed.[1]
Labelling and compute separate
The same study reports that a policy deciding per task uses 33 per cent of training compute across 33 upgrade episodes, with mean quality regret of 0.37 points against always retraining. Relabelling with the old specialist's output retains between 0.96 and 1.05. The authors say this saves annotation work more than compute. The same refresh costs 39.1 GPU-minutes on Banking77 and 255.0 on Spider.[1]
Compilation reduces query-time rereading
Another arXiv preprint targets retrieval-augmented question-answering systems that derive meaning from raw text on every query. The authors propose a queryable layer created at write time. On a held-out sample of 500 broadcast-interview transcripts, compiled claims reached 85.2 per cent correctness from roughly 2,200 reader tokens. The best chunking configuration reached 72.5 per cent from 16,300 tokens. The study reports maintaining the layer is 33.7 times cheaper than rebuilding it and has not been peer reviewed.[2]
Related columns
For more information on this topic, you can read the related columns.