SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing
Organizations: Indigma Innovations · Democritus University of Thrace · Athena Research Center
Abstract
Large language model (LLM) routing aims to select the most suitable model for each incoming query. Most existing routers learn this decision directly from query embeddings, model representations, preference data, or clusters of similar examples. Such approaches can be effective, yet the representation used for routing rarely states what a query actually requires. We introduce SeLMRoute, a routing framework that separates the extraction of candidate-independent semantic evidence from the learning of candidate performance and the application of deployment objectives. A decision model first evaluates a set of interpretable questions about the query, such as its reasoning requirements and use of external knowledge, with each judgment retained as a probability distribution. The resulting probabilistic semantic state is used by a lightweight supervised router to estimate candidate model performance. Routing objectives are applied after performance estimation, which allows the same semantic state to support performance-oriented and cost-aware decisions. On the LLMRouterBench (15 datasets, 20 candidate models, 11,481 queries), SeLMRoute achieves an average accuracy of , while grouped five-fold out-of-fold evaluation reaches , compared with for the strongest fixed candidate. The representation achieves the highest mean performance among the evaluated semantic, dense, lexical, and domain-level representations. In a separate 13-model performance-cost setting, SeLMRoute improves performance in all five grouped splits, with a mean PerfGain of . Our code is available at https://github.com/Indigma-Innovations/SeLMRoute.
Figures & tables
| Probe | Type | Semantic question |
| math_reasoning | Noul | Whether correctness requires mathematical calculation, symbolic manipulation, or quantitative reasoning. |
| code_reasoning | Noul | Whether correctness requires writing, modifying, debugging, or reasoning about executable code. |
| formal_logic | Noul | Whether the task materially requires formal, combinatorial, rule-based, or constraint-satisfaction reasoning. |
| factual_recall | Noul | Whether the request can be solved primarily through factual knowledge with little derivation. |
| social_affective | Noul | Whether correctness depends on emotion, intention, interpersonal meaning, or social context. |
| tool_interaction | Noul | Whether completing the task requires interaction with an external tool, API, software environment, or simulator. |
| Dataset | Queries |
| AIME | 60 |
| BBH | 1,080 |
| EmoryNLP | 697 |
| FinQA | 1,147 |
| GPQA | 198 |
| HumanEval | 164 |
| Method | AvgAcc | Gain@R (%) | Gain@B (%) | Gap@O (%) |
| SeLMRoute ProbabilityMass | ||||
| SeLMRoute Full | ||||
| SeLMRoute Hard | ||||
| GTE-Qwen2 + CatBoost | ||||
| DomainOnly | ||||
| TF-IDF |
| Router | AvgAcc |
| RouterDC | 61.33 |
| GraphRouter | 70.29 |
| EmbedLLM | 71.24 |
| MODEL-SAT | 71.88 |
| Avengers | 71.94 |
| SeLMRoute ProbabilityMass |
| Removed probe | Drop (pp) |
| Constraint density | +0.604 |
| Context integration | +0.511 |
| Current information | +0.487 |
| Tool interaction | +0.314 |
| Formal logic | +0.212 |
| Factual recall | +0.154 |
| Seed | GPT-5 AvgAcc | SeLMRoute AvgAcc | PerfGain | GPT-5 Cost | Cost | Strict CostSave |
| 42 | 65.63 | 67.85 | +3.37% | $125.29 | $133.30 | N/A |
| 3407 | 65.91 | 68.93 | +4.58% | $127.96 | $128.09 | -0.10% |
| 0 | 64.38 | 66.92 | +3.94% | $122.68 | $112.83 | N/A |
| 1 | 64.92 | 65.48 | +0.87% | $127.72 | $108.27 | N/A |
| 2 | 66.78 | 67.12 | +0.52% | $122.23 | $124.16 | -1.58% |
| Mean PerfGain | ||||||
| Config | Mean input | Total input | $/1K queries | Total $ |
| PM-16 | 1,606.7 | 18.45M | 0.0675 | 0.775 |
| Lite-12 | 1,062.7 | 12.20M | 0.0446 | 0.512 |