Existing Large Language Model (LLM) routing methods score LLMs independently to select top-k models. However, this ignores model correlations and enforces a rigid computational budget. Consequently, routers often select redundant models that share failure modes, limiting the overall probability of success. To address this, we propose FlexRouter, a routing framework that explicitly models model complementarity. FlexRouter optimizes for \textit{answer coverage}, maximizing the probability that at least one selected model yields a correct response. This objective aligns with practical inference pipelines where multiple candidate outputs are generated and a downstream verifier or user selects the final one. We formulate routing as a coverage-oriented subset selection problem and model the routing policy using Determinantal Point Processes (DPPs), which naturally capture both model competence and redundancy. To directly optimize coverage without requiring a ground-truth target subset, we introduce a training objective based on marginalizing over failure sets. During inference, we employ a greedy strategy based on marginal log-determinant gains, enabling the router to adaptively determine subset sizes without a predefined budget. Extensive experiments on the large-scale RouterEval benchmark demonstrate that our proposed FlexRouter achieves higher coverage with lower redundancy across both in-domain and out-of-domain tasks than strong baselines while maintaining flexible inference cost.
Figures & tables
Figure 1: Testing Success@10 score of baselines and FlexRouter on in-domain and out-of-domain tasks.
Figure 2: Independent top- k routing often selects highly similar models, leading to redundant outputs and shared failure modes. FlexRouter instead selects complementary models that cover different interpretations, reasoning paths, and implementations, improving answer coverage across tasks.
Figure 3: FlexRouter overview. For a query, the router encoder produces embeddings that define query-dependent competence for each LLM, while learned model embeddings capture pairwise similarity between models. These signals define a DPP kernel whose diagonal entries favor strong models and whose off-diagonals penalize redundant selections, thus encouraging individually strong yet mutually diverse model sets. Greedy selection with early stopping returns an adaptive-size subset.
In-Domain
Out-of-Domain
Method
BBH
MATH
GPQA
IFEval
MuSR
Avg
Ref. score
0.830
0.400
0.397
0.769
0.699
0.619
EmbedLLM
0.9419
0.8264
0.7731
0.9298
0.6931
0.8329
BaRP
0.9124
0.7132
0.8445
0.9187
0.7712
0.8320
Ours
0.9471
0.8075
0.8487
0.9390
0.7738
0.8632
Table 1: Success@10 on the medium-pool setting (3811 models). Ref. score is the reference average accuracy for each task of a representative LLM, such as GPT-4. Best results are bolded ; second-best are underlined .
In-Domain
Out-of-Domain
Method
MMLU
HellaSwag
GSM8K
ARC
TruthfulQA
WinoGrande
Avg
Ref. score
0.864
0.953
0.920
0.852
0.669
0.875
0.855
EmbedLLM
0.9865
0.9776
0.9848
0.9957
0.9682
0.9684
0.9802
BaRP
0.9783
0.7979
0.9432
0.8889
0.8972
0.9795
0.9142
Ours
0.9907
0.9851
0.9886
0.9915
0.9951
0.9976
0.9914
Table 2: Success@10 on the large-pool setting (5000 models). Best results are bolded ; second-best are underlined .
In-Domain
Out-of-Domain
Method
BBH
MATH
GPQA
IFEval
MuSR
Avg
EmbedLLM
0.2114
0.1432
0.2196
0.2524
0.2295
0.2112
BaRP
0.1983
0.1993
0.1990
0.1966
0.1980
0.1982
Ours
0.7249
0.7515
0.7413
0.7895
0.8027
0.7620
Table 3: ILD@10 in the fixed pre-trained embedding space on the medium-pool setting . Higher is better. Best results are bolded .
In-Domain
Out-of-Domain
Method
MMLU
HellaSwag
GSM8K
ARC
TruthfulQA
WinoGrande
Avg
EmbedLLM
0.1302
0.2765
0.1775
0.4536
0.3747
0.2732
0.2810
BaRP
0.5520
0.5521
0.5521
0.5520
0.5521
0.5521
0.5521
Ours
0.4709
0.7052
0.6483
0.7653
0.7396
0.7516
0.6802
Table 4: ILD@10 in the fixed pre-trained embedding space on the large-pool setting . Higher is better. Best results are bolded .
Setting
Kernel
ID Succ
ID ILD
OOD Succ
OOD ILD
Large-pool
RBF
0.9920
0.6019
0.9954
0.6923
Cosine
0.9854
0.6474
0.9962
0.7456
Medium-pool
RBF
0.8804
0.7109
0.8757
0.7883
Cosine
0.9058
0.7392
0.8486
0.7961
Table 5: Kernel design ablation at k=10 . ILD is measured in the fixed embedding space.
Figure 4: Success@10 and Average Subset Size over Stopping Threshold τ . When τ≈0.2 , the router reduces the average subset size substantially while maintaining high coverage.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
τ
Avg Subset Size
Success@10
Size Reduction
Coverage Drop
0.000
10.00
0.9885
0.0 %
0.00 %
0.001
10.00
0.9885
0.0 %
0.00 %
0.010
10.00
0.9885
0.0 %
0.00 %
0.050
10.00
0.9885
0.0 %
0.00 %
0.100
9.95
0.9883
0.5 %
0.02 %
0.200
6.41
0.9772
35.9 %
1.13 %
Appendix
Table 6: Effect of the relative stopping threshold τ on coverage and inference cost (combined ID test, large-pool setting, kmax=10 ). τ is expressed as a fraction of the first marginal log-determinant gain. Size Reduction and Coverage Drop are relative to τ=0 .
Dataset
Category
#Prompts
#LLMs
BBH
Complex Reasoning
5761
3811
MATH
Mathematical Reasoning
1324
3811
GPQA
Graduate-level QA
1192
3811
IFEval
Instruction Following
541
3811
MuSR
Multi-step Reasoning
756
3811
MMLU
Knowledge
14042
5000
Appendix
Table 7: Statistics of datasets used in our work. Each dataset consists of prompts evaluated across a shared pool of candidate LLMs within each setting.
Setting
k
FlexRouter
EmbedLLM
BaRP
Medium-pool
1
0.417
0.542
0.527
Medium-pool
3
0.709
0.707
0.639
Medium-pool
5
0.789
0.765
0.732
Medium-pool
7
0.832
0.805
0.784
Medium-pool
10
0.872
0.840
0.844
Large-pool
1
0.700
0.750
0.665
Appendix
Table 8: Average Success@k across different model-call budgets. FlexRouter is designed for complementary subset construction rather than single-model routing. It becomes stronger once the setting allows multiple models to be selected.
Setting
Method
Avg-Correct@10 ↑
Zero-Correct Rate ↓
Success@10 ↑
Medium-pool
EmbedLLM
5.17
0.160
0.840
Medium-pool
BaRP
4.51
0.156
0.844
Medium-pool
FlexRouter
4.90
0.128
0.872
Large-pool
EmbedLLM
7.94
0.022
0.978
Large-pool
BaRP
6.64
0.084
0.916
Large-pool
FlexRouter
7.49
0.008
0.992
Appendix
Table 9: Pool-composition analysis at k=10 . Avg-Correct@10 measures the average number of correct models in the selected subset, while Zero-Correct Rate measures the fraction of queries for which the selected subset contains no correct model.
Setting
k
FlexRouter
MMR
MaxDiversity
Random-k
Medium-pool
3
0.709
0.706
0.689
0.592
Medium-pool
5
0.789
0.757
0.754
0.697
Medium-pool
10
0.872
0.842
0.837
0.806
Large-pool
3
0.954
0.921
0.915
0.745
Large-pool
5
0.975
0.942
0.935
0.812
Large-pool
10
0.992
0.961
0.950
0.874
Appendix
Table 10: Coverage-oriented baselines across medium- and large-pool settings. FlexRouter outperforms Random-k, MaxDiversity, and MMR across different budgets.
Training Queries
MATH
BBH
GPQA
IFEval
MuSR
Avg Success@10
25%
0.785
0.943
0.836
0.963
0.772
0.860
50%
0.793
0.941
0.845
0.954
0.792
0.865
100%
0.793
0.946
0.878
0.963
0.838
0.884
Appendix
Table 11: Supervision ablation in the medium-pool setting. FlexRouter remains effective when trained with fewer labeled queries.
No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective deployment. Existing routers adopt diverse formulations and implementations, making fair comparison and extension difficult. We present a unified formulation of LLM routing as a sequential decision process characterized by five components: context encoders, model encoders, scoring functions, decision rules, and learning signals, covering single-turn, multi-turn, and personalized routing. Based on this formulation, we develop an automated pipeline for constructing routing supervision and evaluating routers jointly on response quality and inference cost. The resulting benchmark, xRouteBench, spans generic LLM, memory-augmented, vision, time-series, and personalized routing tasks. We further introduce LLMRouter, an open-source modular infrastructure with more than 16 representative routers. Our empirical study shows that learned routers outperform the strongest fixed-model baseline by 14.6% relatively, lightweight routers become more competitive under tight cost constraints, and user-conditioned routing consistently improves personalization.
Tao Feng, Fangxu Yu, Haozhen Zhang +9
University of Illinois Urbana-Champaign · University of Maryland, College Park · 3Nanyang Technological University +2
Large language models (LLMs) achieve impressive performance across multiple domains, but using the most capable model for every query is prohibitive at scale. LLM routing exploits diversity in model capability and cost by assigning each query to a suitable model to balance utility and budget. Current methods have two limitations: (i) they either use heuristics that do not always enforce the budget constraint or impose a fixed per-query budget that cannot adapt across the workload and leads to suboptimal performance; (ii) they require supervised learning on a dense dataset with statistics for every query-model pair, which is expensive to collect. To address these challenges, we formulate LLM routing as a constrained contextual multi-armed bandit problem and introduce WISERouter (WR for short), a framework that supports offline learning from historical interactions as well as online learning with exploration. We further prove that WR-Online achieves a sublinear regret bound of O(T) over a time horizon T. Empirical results on RouterBench and SWE-Bench demonstrate that (i) WR-Offline surpasses existing baselines in performance under a fixed budget and adheres more closely to budget constraints, and (ii) WR-Online achieves comparable performance to the baselines, while using substantially less exploration data.
Large language models are increasingly used in practical systems, making efficient model selection important for reducing deployment cost. LLM routing has emerged as a practical solution for allocating each input query to an appropriate model under a desired cost-performance trade-off. Existing routing methods often estimate model suitability from the surface semantics or embedding similarity of the input query. However, such methods may ignore the underlying difficulty of a query, leading to suboptimal routing decisions. To address the challenge, we propose VDAR-Router, a difficulty-aware retrieval-based routing framework. For each input query, VDAR-Router first generates an explicit difficulty analysis. It then retrieves historical examples with similar difficulty profiles. Based on the retrieved records, it estimates candidate model suitability and selects the model using a reward function that considers both performance and cost. Experiments on three datasets show that VDAR-Router consistently achieves better cost-performance trade-offs than existing baselines. These results demonstrate the effectiveness of difficulty-aware retrieval for training-free LLM routing. Case studies further show that explicit query analysis helps retrieve more relevant examples and supports more reliable routing decisions.
Yu-Chien Tang, Jun-Chen Hung, Wen-Chih Peng +1
National Yang Ming Chiao Tung University · Department of Computer Science Hsinchu, Taiwan