Six model routers fail to beat random selection in an experiment
Researchers testing six commercial model routers found that random selection between two well-chosen models achieved at least comparable results at matched cost. Their English-language experiment links routing decisions to response length as well as question difficulty. The finding comes from one preprint and concerns benchmark tasks, with each response generated once rather than repeated trials on real user traffic.
Artificial Intelligence··Midday
Random selection holds its ground
Six commercial model routers, systems that decide which model receives a query, failed to outperform choosing randomly from a carefully selected pair at the same expenditure in a new preprint. The researchers tested OpenRouter, Microsoft Azure, vLLM, Not Diamond, Nadir and Orca across 14 configurations. Some configurations had a percentage-point gap exceeding 10 against the random baseline.[1]
Short answers change the routing incentive
The cost–accuracy objective can favor sending moderately difficult questions to stronger models: the hardest questions may defeat both cheap and expensive models. Short answers can also earn the same credit for correctness with less generation cost. The authors identify difficulty blindness, preference for escalating short responses, matching a query to its source, and poorly chosen model rosters.[1]
Eight task categories define the experiment
The evaluation spans eight categories drawn from 17 benchmarks, with 800 evaluation questions and answers generated across 26 models. Researchers recorded the chosen model and generated its answer separately under common settings. Questions are in English and each answer was generated once. Their proof-of-concept two-model router achieved only a limited improvement over random selection; the study does not measure repeated responses across real user traffic.[1]