Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing
Authors: Guannan Lai, Gelin Bian, Hao-Xuan Ma, Jun-Peng Jiang, Long Chen, Jian-Dong Liu, Zhi-Hao Tan, Han-Jia Ye
Organizations: School of Artificial Intelligence, Nanjing University · National Key Laboratory for Novel Software Technology, Nanjing University · SinapisAI · The Hong Kong University of Science and Technology
Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on historical queries to collect query--model quality feedback, creating a nontrivial supervision cost before deployment. Existing work largely focuses on serving-time efficiency, overlooking whether the resulting savings are sufficient to recover this upfront expenditure. We further observe that routing quality often saturates well before all query--model feedback is collected, suggesting that dense supervision can be economically over-provisioned. We propose SaveRouter, a sparse-supervision routing framework that selectively acquires informative model feedback and shares capability information across related queries, while retaining query-level refinement for fine-grained routing. We evaluate routing by jointly accounting for supervision expenditure and subsequent serving-time savings. Across four routing benchmarks, the main setting uses only about 33--41% of available training feedback while maintaining competitive or better routing quality, and reduces the break-even deployment volume by approximately 1.9--9.5 times compared with the fastest conventional router. Further analysis shows that acquiring more supervision is not always economically preferable: the supervision level that minimizes serving cost can differ from the one that achieves the earliest payback. Our code is publicly available at https://github.com/LAMDA-Model-Reuse/SaveRouter.
Figures & tables
Figure 1: Dense supervision can be economically over-provisioned. Left: supervision expenditure can substantially delay the break-even point of routers with lower serving-time cost. Right: routing quality often saturates well before the full query–model matrix is observed; EmbedLLM and kNN recover 99% of their fully supervised accuracy with only 30% and 60% supervision, respectively.
Figure 2: Overview of SaveRouter . SaveRouter first groups related training queries and adaptively acquires sparse query–model feedback using capability and uncertainty. It then estimates model capability through a structured group–model prior, evidence-based correction from the acquired observations, and a query-specific residual predictor. At deployment time, the resulting quality estimates are combined with estimated serving costs for cost-aware routing.
LLMRouterBench
Mixinstruct
Method
Ps↑
CR ↓
SA-BEP ↓
SA-CR@1M ↓
Ps↑
CR ↓
SA-BEP ↓
SA-CR@1M ↓
EmbedLLM
0.6094
0.5246
15.6K
0.5321
0.7488
0.9924
28.60M
1.2088
kNN
0.6230
0.3770
11.9K
0.3845
0.7484
0.9814
11.63M
1.1977
OmniRouter
0.6218
0.3564
11.5K
0.3638
0.7435
∞
∞
∞
RMSoftmax
0.6126
0.6328
20.2K
0.6403
0.7495
0.9893
20.19M
1.2056
TRouter
0.6145
0.5221
15.6K
0.5295
0.7488
1.0000
∞
∞
Table 1: Main results on four routing benchmarks. Higher Ps is better; lower CR, SA-BEP, and SA-CR@1M are better. ∞ denotes an unreachable target operating point; for supervision-amortized metrics, it also indicates non-positive per-query saving. For sparse- or partial-feedback methods, every model invocation used to acquire feedback is charged separately.
Figure 4
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Symbol
Value
Global prior parameter
α0
1
Global prior parameter
β0
1
Acquisition prior strength
τ0
10
UCB exploration coefficient
βucb
0.35
Additive-prior Ridge weight
λprior
10
Final shrinkage strength
τ
40
Appendix
Table A1: Core hyperparameters of SaveRouter .
Benchmark
Training groups
New-query assignment
LLMRouterBench
10 task groups
Logistic classifier
Mixinstruct
Single global group
Single group
MMRBench
7 dataset groups
Logistic classifier
RouterBench
85 task groups
Logistic classifier
Appendix
Table A2: Query grouping used in the main experiments.
Benchmark
#Queries
#Train
#Test
#Models
#Groups
Train Pairs
LLMRouterBench
12,446
2,489
9,957
12
10
29,868 / 29,868
Mixinstruct
110,000
22,000
88,000
12
1
264,000 / 264,000
MMRBench
10,370
2,074
8,296
10
7
20,149 / 20,740
RouterBench
36,497
7,299
29,198
11
85
80,289 / 80,289
Appendix
Table A3: Statistics of the routing benchmarks used in our experiments. Available train pairs count query–model pairs for which both quality and cost fields are available.
Method
Configuration
EmbedLLM
α=0.001 ; 30 epochs; learning rate 10−3 .
kNN
k=300 .
OmniRouter
top- k=50 ; γ=δ=0.6 ; 50 epochs.
RMSoftmax
30 linearly spaced cost weights from 0 to 1000 ; 300 epochs.
TRouter
hidden dimension 256; dropout 0.1; temperature 0.07; 20 epochs.
UniRoute
10 clusters; mapping network trained for 5 epochs.
Appendix
Table A4: Key configurations of the compared routing methods.
Benchmark
K
Sup. (%)
Ps↑
CR ↓
C0↓
SA-BEP ↓
SA-CR@1M ↓
LLMRouterBench
1
8.33
0.621874
0.2395
44.64
1,311
0.2405
2
16.67
0.631415
0.2357
85.21
2,489
0.2376
3
25.00
0.630160
0.2228
146.69
4,213
0.2261
4
33.33
0.633825
0.2320
181.09
5.3K
0.2360
6
50.00
0.627699
0.2180
241.15
6,884
0.2234
8
66.67
0.627147
0.2260
279.86
8,071
0.2322
Appendix
Table A5: Effect of supervision budget K across four routing benchmarks. Higher Ps is better; lower CR, C0 , SA-BEP, and SA-CR@1M are better. “–” indicates that the router does not reach the best-single-model quality target, so the corresponding supervision-amortized metrics are undefined.
Method
Ps↑
CR ↓
SA-BEP ↓
SA-CR@1M ↓
Dense-supervision routers
EmbedLLM
0.776135
∞
∞
∞
kNN
0.795418
0.816906
4,772,514
1.6907
MLP
0.776705
∞
∞
∞
SVM
0.718650
∞
∞
∞
RouteLLM-MF
0.780365
∞
∞
∞
Appendix
Table A6: Results on RouterEval. SaveRouter uses K=108 . Higher Ps is better; lower CR, SA-BEP, and SA-CR@1M are better. ∞ denotes an unreachable target operating point.
Train/Test
Ps↑
CR ↓
SaveRouter
Best Dense
ΔPs (pp)
SaveRouter
Best Dense
Δ CR
20/80
0.6338
0.6230
+1.08
0.2320
0.3564
34.9%
40/60
0.634842
0.624665
+1.018
0.1888
0.3275
42.4%
60/40
0.633461
0.626029
+0.743
0.2137
0.2965
27.9%
80/20
0.641566
0.636145
+0.542
0.1412
0.2275
38.0%
Appendix
Table A7: Sensitivity to the training-set fraction on LLMRouterBench. “Best dense” denotes the strongest dense baseline for the corresponding metric among EmbedLLM, kNN, OmniRouter, TRouter, and InferenceDynamics. ΔPs is the absolute improvement of SaveRouter in percentage points, and Δ CR is the relative reduction in CR.
Configuration
Ps↑
CR ↓
SA-BEP ↓
SA-CR@1M ↓
Main: γ=2,τ=40,λctx=200
0.633825
0.2320
5,263
0.2360
λctx=100
0.631365
0.2182
5,171
0.2223
λctx=50
0.625791
0.2208
5,188
0.2248
γ=1
0.633373
0.2517
5,402
0.2558
γ=4
0.624385
0.2430
5,340
0.2471
τ=20
0.628955
0.2024
5,068
0.2065
Appendix
Table A8: Hyperparameter robustness on LLMRouterBench with fixed K=4 supervision. Higher Ps is better; lower CR, SA-BEP, and SA-CR@1M are better.
Figure A1: Full quality–cost frontiers across four routing benchmarks. Each solid curve shows all evaluated SaveRouter operating points, while markers denote the peak-quality endpoint of each baseline method. The horizontal axis reports serving cost normalized by the best single model within each benchmark; lower is better, while higher quality is better.
Method
0.90Qb
0.925Qb
0.95Qb
0.975Qb
1.00Qb
EmbedLLM
0.0286 / 7.7K / 0.0360
0.0953 / 8.2K / 0.1027
0.5246 / 15.6K / 0.5321
0.5246 / 15.6K / 0.5321
0.5246 / 15.6K / 0.5321
kNN
0.0863 / 8.1K / 0.0938
0.1107 / 8.4K / 0.1182
0.1530 / 8.8K / 0.1605
0.2490 / 9.9K / 0.2564
0.3770 / 11.9K / 0.3845
OmniRouter
0.1419 / 8.7K / 0.1493
0.1618 / 8.9K / 0.1693
0.2037 / 9.3K / 0.2112
0.2579 / 10.0K / 0.2653
0.3564 / 11.5K / 0.3638
RMSoftmax
0.0285 / 7.7K / 0.0359
0.6328 / 20.2K / 0.6403
0.6328 / 20.2K / 0.6403
0.6328 / 20.2K / 0.6403
0.6328 / 20.2K / 0.6403
TRouter
0.0328 / 7.7K / 0.0402
0.1500 / 8.7K / 0.1575
0.2447 / 9.8K / 0.2521
0.2995 / 10.6K / 0.3069
0.5221 / 15.6K / 0.5295
UniRoute
0.0285 / 7.7K / 0.0359
0.1962 / 9.2K / 0.2036
0.5583 / 16.8K / 0.5657
0.5583 / 16.8K / 0.5657
0.5583 / 16.8K / 0.5657
Appendix
Table A9: Performance across quality targets on LLMRouterBench. Each entry reports CR / SA-BEP / SA-CR@1M; lower values are better.
Method
0.90Qb
0.925Qb
0.95Qb
0.975Qb
1.00Qb
EmbedLLM
0.5000 / 432.7K / 0.7163
0.5000 / 432.7K / 0.7163
0.9924 / 28.60M / 1.2088
0.9924 / 28.60M / 1.2088
0.9924 / 28.60M / 1.2088
kNN
0.5000 / 432.7K / 0.7163
0.5000 / 432.7K / 0.7163
0.8561 / 1.50M / 1.0725
0.8561 / 1.50M / 1.0725
0.9814 / 11.63M / 1.1977
OmniRouter
0.5596 / 491.2K / 0.7759
0.5596 / 491.2K / 0.7759
0.6292 / 583.4K / 0.8455
0.8097 / 1.14M / 1.0261
∞ / ∞ / ∞
RMSoftmax
0.5000 / 432.7K / 0.7163
0.5000 / 432.7K / 0.7163
0.9893 / 20.19M / 1.2056
0.9893 / 20.19M / 1.2056
0.9893 / 20.19M / 1.2056
TRouter
0.5000 / 432.7K / 0.7163
0.5000 / 432.7K / 0.7163
0.6441 / 607.9K / 0.8604
0.9879 / 17.95M / 1.2043
1.0000 / ∞ / ∞
UniRoute
0.5000 / 432.7K / 0.7163
0.5000 / 432.7K / 0.7163
1.0000 / ∞ / ∞
1.0000 / ∞ / ∞
1.0000 / ∞ / ∞
Appendix
Table A10: Performance across quality targets on Mixinstruct. Each entry reports CR / SA-BEP / SA-CR@1M; lower values are better.
Method
0.90Qb
0.925Qb
0.95Qb
0.975Qb
1.00Qb
EmbedLLM
1.0000 / ∞ / ∞
1.0000 / ∞ / ∞
1.0000 / ∞ / ∞
1.0000 / ∞ / ∞
1.0000 / ∞ / ∞
kNN
0.5503 / 12.2K / 0.5557
0.7515 / 22.0K / 0.7569
0.9274 / 75.3K / 0.9329
1.0000 / ∞ / ∞
1.0000 / ∞ / ∞
OmniRouter
0.2911 / 7.7K / 0.2965
0.4700 / 10.3K / 0.4755
0.6581 / 16.0K / 0.6635
0.8215 / 30.6K / 0.8269
∞ / ∞ / ∞
RMSoftmax
1.0350 / ∞ / ∞
1.0350 / ∞ / ∞
1.0350 / ∞ / ∞
1.0350 / ∞ / ∞
1.0350 / ∞ / ∞
TRouter
0.2103 / 6.9K / 0.2158
0.6730 / 16.7K / 0.6784
0.6730 / 16.7K / 0.6784
0.6730 / 16.7K / 0.6784
0.8396 / 34.1K / 0.8451
UniRoute
0.8173 / 29.9K / 0.8227
0.8173 / 29.9K / 0.8227
0.8173 / 29.9K / 0.8227
0.8173 / 29.9K / 0.8227
∞ / ∞ / ∞
Appendix
Table A11: Performance across quality targets on MMRBench. Each entry reports CR / SA-BEP / SA-CR@1M; lower values are better.
Method
0.90Qb
0.925Qb
0.95Qb
0.975Qb
1.00Qb
EmbedLLM
0.0742 / 24.8K / 0.0972
0.2648 / 31.3K / 0.2878
0.9296 / 326.4K / 0.9526
0.9296 / 326.4K / 0.9526
0.9296 / 326.4K / 0.9526
kNN
0.0655 / 24.6K / 0.0885
0.1443 / 26.9K / 0.1672
0.4101 / 39.0K / 0.4331
0.7455 / 90.3K / 0.7684
∞ / ∞ / ∞
OmniRouter
0.1617 / 27.4K / 0.1847
0.1811 / 28.1K / 0.2040
0.3522 / 35.5K / 0.3752
0.6616 / 67.9K / 0.6846
∞ / ∞ / ∞
RMSoftmax
0.9865 / 1.70M / 1.0095
0.9865 / 1.70M / 1.0095
0.9865 / 1.70M / 1.0095
0.9865 / 1.70M / 1.0095
0.9865 / 1.70M / 1.0095
TRouter
0.0901 / 25.3K / 0.1131
0.3151 / 33.5K / 0.3381
0.4450 / 41.4K / 0.4680
0.5987 / 57.3K / 0.6217
∞ / ∞ / ∞
UniRoute
0.0741 / 24.8K / 0.0970
0.9327 / 341.3K / 0.9557
0.9327 / 341.3K / 0.9557
0.9327 / 341.3K / 0.9557
0.9327 / 341.3K / 0.9557
Appendix
Table A12: Performance across quality targets on RouterBench. Each entry reports CR / SA-BEP / SA-CR@1M; lower values are better.
Ps↑
CR ↓
SA-BEP ↓
SA-CR@1M ↓
Benchmark
Cached
Fresh
Cached
Fresh
LLMRouterBench
0.6181
0.3161
10.8K
114.0K
0.3235
0.3941
Mixinstruct
0.7482
0.9695
6.13M
44.66M
1.1565
2.3329
MMRBench
0.7448
0.9052
57.6K
640.8K
0.9107
0.9660
RouterBench
0.8053
0.9355
329.2K
3.25M
0.9567
1.1449
Appendix
Table A13: Sensitivity of BaRP to fresh versus cached feedback accounting. Ps and CR are unchanged because caching affects only supervision expenditure. Lower SA-BEP and SA-CR@1M are better.
Benchmark
Arrival
∣Ω+∣
ρC,+
Ps,10
Ps,100
Ps,dense
CR10
CR100
CRdense
SA-BEP 10
SA-CR@1M 10
LLMRouterBench
Cheap
254
11.92%
0.6276
0.6273
0.6230
0.2296
0.2261
0.3564
4,921
0.2334
Median
254
9.35%
0.6263
0.6340
0.6230
0.2650
0.2287
0.3564
5,610
0.2691
Expensive
254
10.97%
0.6339
0.6305
0.6230
0.2606
0.2412
0.3564
3,887
0.2635
Mixinstruct
Cheap
2,200
10.00%
0.7486
0.7496
0.7495
0.9847
0.9516
0.9814
6,088,547
1.0779
Median
2,200
10.00%
0.7498
0.7498
0.7495
0.9354
0.9354
0.9814
1,257,581
1.0166
Expensive
2,200
10.00%
0.7498
0.7497
0.7495
0.9368
0.9569
0.9814
1,287,017
1.0181
Appendix
Table A14: Sparse model-pool expansion with K=4 supervision for the existing model pool. Subscripts 10 and 100 denote using approximately 10% and 100% of the available training feedback for the arriving model. ρC,+ is the acquisition cost of the 10% feedback relative to fully labeling the arriving model. Dense Ps and CR are independent per-metric optima over the fully supervised baselines and need not correspond to the same router. Higher Ps is better; lower ρC,+ , CR, SA-BEP, and SA-CR@1M are better.
No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective deployment. Existing routers adopt diverse formulations and implementations, making fair comparison and extension difficult. We present a unified formulation of LLM routing as a sequential decision process characterized by five components: context encoders, model encoders, scoring functions, decision rules, and learning signals, covering single-turn, multi-turn, and personalized routing. Based on this formulation, we develop an automated pipeline for constructing routing supervision and evaluating routers jointly on response quality and inference cost. The resulting benchmark, xRouteBench, spans generic LLM, memory-augmented, vision, time-series, and personalized routing tasks. We further introduce LLMRouter, an open-source modular infrastructure with more than 16 representative routers. Our empirical study shows that learned routers outperform the strongest fixed-model baseline by 14.6% relatively, lightweight routers become more competitive under tight cost constraints, and user-conditioned routing consistently improves personalization.
Tao Feng, Fangxu Yu, Haozhen Zhang +9
University of Illinois Urbana-Champaign · University of Maryland, College Park · 3Nanyang Technological University +2
Inference costs for large language model (LLM) applications are rapidly growing, driven by surging demand and rising infrastructure cost. Users expect high-quality responses, and in commercial settings this is formally codified in Service Level Agreements (SLAs), creating a fundamental tension between cost and quality. Recent progress on cost-aware LLM request routing has shown potential to resolve this tension, but existing approaches rely on complete feedback signals, offline training, extensive per-workload tuning, and most lack SLA guarantees or inference-time adaptivity. We introduce SLARouter, an online routing algorithm that learns a cost-optimal policy from the sparse, one-sided user feedback available in production systems. SLARouter provides theoretical guarantees for both cost optimality and strict SLA compliance. Experiments across a wide range of LLM benchmarks show that SLARouter satisfies SLA constraints without the need for per-benchmark tuning, reducing operating cost by up to 2.2x over existing baselines.
Herbert Woisetschläger, Arastun Mammadli, Ryan Zhang +1
Technical University of Munich Germany · University of Exeter United Kingdom · Horace Greeley High School United States
Large language models (LLMs) achieve impressive performance across multiple domains, but using the most capable model for every query is prohibitive at scale. LLM routing exploits diversity in model capability and cost by assigning each query to a suitable model to balance utility and budget. Current methods have two limitations: (i) they either use heuristics that do not always enforce the budget constraint or impose a fixed per-query budget that cannot adapt across the workload and leads to suboptimal performance; (ii) they require supervised learning on a dense dataset with statistics for every query-model pair, which is expensive to collect. To address these challenges, we formulate LLM routing as a constrained contextual multi-armed bandit problem and introduce WISERouter (WR for short), a framework that supports offline learning from historical interactions as well as online learning with exploration. We further prove that WR-Online achieves a sublinear regret bound of O(T) over a time horizon T. Empirical results on RouterBench and SWE-Bench demonstrate that (i) WR-Offline surpasses existing baselines in performance under a fixed budget and adheres more closely to budget constraints, and (ii) WR-Online achieves comparable performance to the baselines, while using substantially less exploration data.