FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing
Organizations: Virginia Tech · Independent Researcher · University of Southern California · Adobe Research · Dolby Labs
Abstract
Existing Large Language Model (LLM) routing methods score LLMs independently to select top- models. However, this ignores model correlations and enforces a rigid computational budget. Consequently, routers often select redundant models that share failure modes, limiting the overall probability of success. To address this, we propose FlexRouter, a routing framework that explicitly models model complementarity. FlexRouter optimizes for \textit{answer coverage}, maximizing the probability that at least one selected model yields a correct response. This objective aligns with practical inference pipelines where multiple candidate outputs are generated and a downstream verifier or user selects the final one. We formulate routing as a coverage-oriented subset selection problem and model the routing policy using Determinantal Point Processes (DPPs), which naturally capture both model competence and redundancy. To directly optimize coverage without requiring a ground-truth target subset, we introduce a training objective based on marginalizing over failure sets. During inference, we employ a greedy strategy based on marginal log-determinant gains, enabling the router to adaptively determine subset sizes without a predefined budget. Extensive experiments on the large-scale RouterEval benchmark demonstrate that our proposed FlexRouter achieves higher coverage with lower redundancy across both in-domain and out-of-domain tasks than strong baselines while maintaining flexible inference cost.
Figures & tables
| In-Domain | Out-of-Domain | |||||
|---|---|---|---|---|---|---|
| Method | BBH | MATH | GPQA | IFEval | MuSR | Avg |
| Ref. score | 0.830 | 0.400 | 0.397 | 0.769 | 0.699 | 0.619 |
| EmbedLLM | 0.9419 | 0.8264 | 0.7731 | 0.9298 | 0.6931 | 0.8329 |
| BaRP | 0.9124 | 0.7132 | 0.8445 | 0.9187 | 0.7712 | 0.8320 |
| Ours | 0.9471 | 0.8075 | 0.8487 | 0.9390 | 0.7738 | 0.8632 |
| In-Domain | Out-of-Domain | ||||||
|---|---|---|---|---|---|---|---|
| Method | MMLU | HellaSwag | GSM8K | ARC | TruthfulQA | WinoGrande | Avg |
| Ref. score | 0.864 | 0.953 | 0.920 | 0.852 | 0.669 | 0.875 | 0.855 |
| EmbedLLM | 0.9865 | 0.9776 | 0.9848 | 0.9957 | 0.9682 | 0.9684 | 0.9802 |
| BaRP | 0.9783 | 0.7979 | 0.9432 | 0.8889 | 0.8972 | 0.9795 | 0.9142 |
| Ours | 0.9907 | 0.9851 | 0.9886 | 0.9915 | 0.9951 | 0.9976 | 0.9914 |
| In-Domain | Out-of-Domain | |||||
|---|---|---|---|---|---|---|
| Method | BBH | MATH | GPQA | IFEval | MuSR | Avg |
| EmbedLLM | 0.2114 | 0.1432 | 0.2196 | 0.2524 | 0.2295 | 0.2112 |
| BaRP | 0.1983 | 0.1993 | 0.1990 | 0.1966 | 0.1980 | 0.1982 |
| Ours | 0.7249 | 0.7515 | 0.7413 | 0.7895 | 0.8027 | 0.7620 |
| In-Domain | Out-of-Domain | ||||||
|---|---|---|---|---|---|---|---|
| Method | MMLU | HellaSwag | GSM8K | ARC | TruthfulQA | WinoGrande | Avg |
| EmbedLLM | 0.1302 | 0.2765 | 0.1775 | 0.4536 | 0.3747 | 0.2732 | 0.2810 |
| BaRP | 0.5520 | 0.5521 | 0.5521 | 0.5520 | 0.5521 | 0.5521 | 0.5521 |
| Ours | 0.4709 | 0.7052 | 0.6483 | 0.7653 | 0.7396 | 0.7516 | 0.6802 |
| Setting | Kernel | ID Succ | ID ILD | OOD Succ | OOD ILD |
|---|---|---|---|---|---|
| Large-pool | RBF | 0.9920 | 0.6019 | 0.9954 | 0.6923 |
| Cosine | 0.9854 | 0.6474 | 0.9962 | 0.7456 | |
| Medium-pool | RBF | 0.8804 | 0.7109 | 0.8757 | 0.7883 |
| Cosine | 0.9058 | 0.7392 | 0.8486 | 0.7961 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Avg Subset Size | Success@10 | Size Reduction | Coverage Drop | |
|---|---|---|---|---|
| 0.000 | 10.00 | 0.9885 | 0.0 % | 0.00 % |
| 0.001 | 10.00 | 0.9885 | 0.0 % | 0.00 % |
| 0.010 | 10.00 | 0.9885 | 0.0 % | 0.00 % |
| 0.050 | 10.00 | 0.9885 | 0.0 % | 0.00 % |
| 0.100 | 9.95 | 0.9883 | 0.5 % | 0.02 % |
| 0.200 | 6.41 | 0.9772 | 35.9 % | 1.13 % |
| Dataset | Category | #Prompts | #LLMs |
| BBH | Complex Reasoning | 5761 | 3811 |
| MATH | Mathematical Reasoning | 1324 | 3811 |
| GPQA | Graduate-level QA | 1192 | 3811 |
| IFEval | Instruction Following | 541 | 3811 |
| MuSR | Multi-step Reasoning | 756 | 3811 |
| MMLU | Knowledge | 14042 | 5000 |
| Setting | FlexRouter | EmbedLLM | BaRP | |
|---|---|---|---|---|
| Medium-pool | 1 | 0.417 | 0.542 | 0.527 |
| Medium-pool | 3 | 0.709 | 0.707 | 0.639 |
| Medium-pool | 5 | 0.789 | 0.765 | 0.732 |
| Medium-pool | 7 | 0.832 | 0.805 | 0.784 |
| Medium-pool | 10 | 0.872 | 0.840 | 0.844 |
| Large-pool | 1 | 0.700 | 0.750 | 0.665 |
| Setting | Method | Avg-Correct@10 | Zero-Correct Rate | Success@10 |
|---|---|---|---|---|
| Medium-pool | EmbedLLM | 5.17 | 0.160 | 0.840 |
| Medium-pool | BaRP | 4.51 | 0.156 | 0.844 |
| Medium-pool | FlexRouter | 4.90 | 0.128 | 0.872 |
| Large-pool | EmbedLLM | 7.94 | 0.022 | 0.978 |
| Large-pool | BaRP | 6.64 | 0.084 | 0.916 |
| Large-pool | FlexRouter | 7.49 | 0.008 | 0.992 |
| Setting | FlexRouter | MMR | MaxDiversity | Random-k | |
|---|---|---|---|---|---|
| Medium-pool | 3 | 0.709 | 0.706 | 0.689 | 0.592 |
| Medium-pool | 5 | 0.789 | 0.757 | 0.754 | 0.697 |
| Medium-pool | 10 | 0.872 | 0.842 | 0.837 | 0.806 |
| Large-pool | 3 | 0.954 | 0.921 | 0.915 | 0.745 |
| Large-pool | 5 | 0.975 | 0.942 | 0.935 | 0.812 |
| Large-pool | 10 | 0.992 | 0.961 | 0.950 | 0.874 |
| Training Queries | MATH | BBH | GPQA | IFEval | MuSR | Avg Success@10 |
|---|---|---|---|---|---|---|
| 25% | 0.785 | 0.943 | 0.836 | 0.963 | 0.772 | 0.860 |
| 50% | 0.793 | 0.941 | 0.845 | 0.954 | 0.792 | 0.865 |
| 100% | 0.793 | 0.946 | 0.878 | 0.963 | 0.838 | 0.884 |