Eigen RadarAI
Analysis

A sampling correction brings combined language-model policies closer to their target

A single preprint examines a mismatch between combining token probabilities and the intended distribution over complete answers. Its theorem places a Metropolis–Hastings correction at least as close to that target as importance resampling with the same candidate budget. Mathematics and conversation experiments accompany the analysis, while distributional fidelity remains separate from task success and general safety.

Artificial Intelligence··Evening
A hand lifts a fourth sheet beside three separate sheets on a computing desk; the blank reverse sides face the viewer.

Combining tokens can miss the intended answer distribution

A language model can have separate policies trained for goals such as accuracy and brevity. Combining them during generation allows a trade-off without training again for every preference. A single preprint examines why combining their next-token probabilities can miss the intended distribution over complete answers. The local calculation does not include the total probability of future continuations, so a sequence can receive a different weight from the one intended.[1]

The theorem compares equal candidate budgets

The authors analyze an existing Metropolis–Hastings correction using normalization values recorded as responses are generated. They compare it with importance resampling, which selects from an independently generated candidate pool. At the same candidate budget, the theorem places the corrected distribution no farther from the target under the covered family of convex divergence measures. Selection occurs after generation and does not require an additional reward model to score each candidate.[1]

A shorter chosen answer can still require more generation

Experiments range from small problems with enumerable answers to full responses from Gemma-4-E2B-it accuracy and brevity adapters. The larger tests use GSM8K mathematics and HH-RLHF conversations; the latter are scored through reward-model proxies. Every candidate still has to be generated, even if the selected answer is short. Neither sampler can recover an answer absent from its pool. Closeness to the specified distribution therefore remains distinct from general safety or success on every task.[1]

References

  1. News sourcearXivMetropolis–Hastings analysis tightens sampling for composed LLM policies↩1↩2↩3