cs.AISep 29, 2026

Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing

Authors: Guannan Lai, Gelin Bian, Hao-Xuan Ma, Jun-Peng Jiang, Long Chen, Jian-Dong Liu, Zhi-Hao Tan, Han-Jia Ye

Organizations: School of Artificial Intelligence, Nanjing University · National Key Laboratory for Novel Software Technology, Nanjing University · SinapisAI · The Hong Kong University of Science and Technology

Abstract

Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on historical queries to collect query--model quality feedback, creating a nontrivial supervision cost before deployment. Existing work largely focuses on serving-time efficiency, overlooking whether the resulting savings are sufficient to recover this upfront expenditure. We further observe that routing quality often saturates well before all query--model feedback is collected, suggesting that dense supervision can be economically over-provisioned. We propose SaveRouter, a sparse-supervision routing framework that selectively acquires informative model feedback and shares capability information across related queries, while retaining query-level refinement for fine-grained routing. We evaluate routing by jointly accounting for supervision expenditure and subsequent serving-time savings. Across four routing benchmarks, the main setting uses only about 33--41% of available training feedback while maintaining competitive or better routing quality, and reduces the break-even deployment volume by approximately 1.9--9.5 times compared with the fastest conventional router. Further analysis shows that acquiring more supervision is not always economically preferable: the supervision level that minimizes serving cost can differ from the one that achieves the earliest payback. Our code is publicly available at https://github.com/LAMDA-Model-Reuse/SaveRouter.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

    Aug 7, 2026Tao Feng, Fangxu Yu, Haozhen Zhang +9Large Language Model RoutingInference Cost

  2. Cost-Optimal LLM Routing with Limited User Feedback under User Satisfaction Guarantees

    Jun 12, 2026Herbert Woisetschläger, Arastun Mammadli, Ryan Zhang +1Large Language Model RoutingLarge Language Model Benchmarks

  3. WISERouter: LLM Routing with Workload Budget Constraint

    Jul 26, 2026Yifei Li, Zihui Gao, Laks V. S. LakshmananLarge Language Model RoutingWorkload