cs.CLOct 5, 2026

Breaking the Tie: A Cluster-Aware Routing Framework for Large Language Models

Authors: Yao Lu, Zhaiyuan Ji, Yaxin Gao, Zeyu Wang, Zhe Tang, Jiaheng Wei, Zhaowei Zhu, Shanqing Yu, +1 more

Organizations: Institute of Cyberspace Security, Zhejiang University of Technology · Binjiang Institute of Artificial Intelligence, Zhejiang University of Technology · Hong Kong University of Science and Technology (Guangzhou)

Abstract

With the rapid development of artificial intelligence, the emergence of various Large Language Models (LLMs) has created a rich model ecosystem. However, this also brings a key challenge: how to select the optimal model for a specific user query. LLM routing addresses this need by dynamically assigning queries to the most suitable expert in the pool of candidate models. However, existing routing frameworks often simplify this process to a standard classification task; thus, a critical vulnerability is exposed when multiple candidate models correctly answer the same query. We formalize this capability overlap as routing noise, which misleads the router with arbitrarily correct candidate models, ultimately leading to routing collapse (a severe decline in generalization ability on unseen tasks). To address this problem, we propose a novel Cluster-Aware Soft-Labeling Routing (CASLR) framework. CASLR shifts the evaluation paradigm from the success of a single query to macro-domain consensus by replacing traditional one-hot vectors with a masked softmax mechanism. Specifically, for experts who answer incorrectly, we penalize their target probability to zero; for the remaining candidates, we directly compute continuous fine-grained soft labels based on their global clustering utility scores. We then use these refined soft labels to supervise a lightweight router. Specifically, the framework not only demonstrates superior accuracy on multiple benchmarks, but also outperforms Llama-3.3-70B-Instruct by 7.80% in overall average performance. Furthermore, the extremely low routing inference latency of only 1.13s further confirms that CASLR can achieve efficient system scheduling with almost zero additional overhead, while ensuring high response quality.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing

    Sep 28, 2026Vasilis Perifanis, Nikolaos Pavlidis, Symeon SymeonidisLLM Routing

  2. LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

    Aug 7, 2026Tao Feng, Fangxu Yu, Haozhen Zhang +9LLM EvaluationLLM Routing

  3. VDAR-Router: Adaptive LLMs Routing via Verbalized Query Difficulty Analysis Retrieval

    Jul 20, 2026Yu-Chien Tang, Jun-Chen Hung, Wen-Chih Peng +1LLM Inference EfficiencyAdaptive Model Routing