LLM Routing

LLM: Large Language Model

Momentum

30 papers in the last four weeks, up 88% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 224

All topics
CardsList
  1. Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System

    Oct 6, 2026Panagiotis Kasnesis, Christos Chatzigeorgiou, Lazaros Toumanidis +1Small Language ModelsLLM Agent Evaluation

  2. Do I Need the Cloud? Uncertainty-Aware Step-Level Handoff for Small Language Model Agents

    Oct 6, 2026Abolfazl YounesiSmall Language ModelsLLM Agents

  3. Cross-lingual Calibration of Pre-Generation Success Probes for Multilingual LLM Routing

    Oct 5, 2026Andrea Paganelli, Stefano Civelli, Pietro Bernardelle +1Confidence Estimation in Language ModelsLanguage Model Calibration

  4. Breaking the Tie: A Cluster-Aware Routing Framework for Large Language Models

    Oct 5, 2026Yao Lu, Zhaiyuan Ji, Yaxin Gao +6LLM Routing

  5. OR for AI That Does OR: Routing LLMs up the Escalator inside the OSCAR Framework

    Oct 1, 2026Jinzhi Bu, Haixin Tang, Huanan ZhangLarge Language Model-Guided OptimizationOptimal Stopping

  6. Adaptive-GEPA: Make Your Harness Fit Heterogeneous Requests

    Sep 30, 2026Tianyu Chen, Yasi Zhang, Ruiyi Wang +3Large Language Model-Guided OptimizationLLM-Based Program Synthesis

  7. FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing

    Sep 29, 2026Wang Wei, Harry Yang, Tiankai Yang +6Adaptive Model RoutingLLM Routing

  8. You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference

    Sep 29, 2026Liang He, Jingbo Wen, Yixiong Chen +4LLM ServingCost-Aware Inference

  9. Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models

    Sep 29, 2026Yury Nahshan, Nati Daniel, Jacob Goldberger +1Mixture-of-Experts Language ModelsParameter-Free Mixture-of-Experts Routing

  10. Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing

    Sep 29, 2026Guannan Lai, Gelin Bian, Hao-Xuan Ma +5LLM ServingCost-Aware Inference

  11. Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing

    Sep 29, 2026Guannan Lai, Han-Jia YeLLM Routing

  12. SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing

    Sep 28, 2026Vasilis Perifanis, Nikolaos Pavlidis, Symeon SymeonidisLLM Routing

  13. RSI-Router: Evolving Subtask-Level LLM Routing and Skills for Cost-Efficient Agents

    Sep 28, 2026Hao Li, Hangfan Zhang, Zhiyao Cui +5Cost-Aware InferenceLLM Agent Workflow Optimization

  14. A Persistent State for Auditable Mixture-of-Experts Routing

    Sep 28, 2026Abdurrahman Javat, Allan KazakovLLM AuditingState Tracking

  15. FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents

    Sep 28, 2026Xi Xiao, Yunbei Zhang, Chen Liu +7LLM GroundingLLM Inference Scheduling

  16. MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference

    Sep 28, 2026Junfeng Wu, Zehao Fan, Hadjer Benmeziane +3Expert OffloadingMixture-of-Experts Language Models

  17. Learning the Cost of Reliable Inference

    Sep 23, 2026Dimitrios Rontogiannis, Ander Artola Velasco, Manuel Gomez RodriguezCost-Aware InferenceLLM Routing

  18. PentestChain: A Cost-Aware, MCP-Orchestrated Framework for Automated Penetration Testing with Free-Tier LLMs

    Sep 16, 2026Rushabh Vipulkumar Patel, Dipo Dunsin, Mohammed Almaiah +1Automated Penetration TestingLLMs for Cybersecurity

  19. Beyond the Previous Layer: Residual Predictive Structure in Sparse MoE Routing

    Sep 15, 2026Hao Li, Yasuyuki Tahara, Yuichi SeiSparse Mixture-of-ExpertsExpert Routing

  20. One Size Does Not Fit All! Dynamic Retriever and Generator Selection for RAG

    Sep 15, 2026Neeraj Anand, Payel Santra, Partha Basuchowdhuri +2Retrieval-Augmented GenerationAdaptive RAG

  21. Quality-Constrained Routing over a Fixed Pool of Quantized Mixture-of-Experts Instances

    Sep 14, 2026Zhenghong Huang, Hongfan Wu, Jiheng ZhangMixture-of-Experts InferenceMixture-of-Experts Quantization

  22. Route, Don't Fix: Regime-Dependent Decoding Correction and a Trajectory-Gated Router for Reliable Clinical LLM Answer Selection

    Sep 13, 2026Zeyu Dong, Benjamin Wang, Joyee W. JinAdaptive InferenceLanguage Model Decoding

  23. SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model Conversations

    Sep 12, 2026Yu Wang, Yuchen Li, Rui Kong +11Multi-Turn Dialogue EvaluationLLM Routing

  24. Router Prior Bias: Preserving Base Routing Structure in MoE Post-Training

    Sep 8, 2026Jaedeok Lee, Keonwoo Kim, Dongyoon Han +3LLM RoutingMixture-of-Experts Models

  25. Signed Rescue Routing: Harm-Aware Cascades for Efficient LLM Inference

    Sep 7, 2026Zheyuan Wang, Siyu Li, Peiqiao Song +3LLM Inference EfficiencyLLM Routing

  26. SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology

    Sep 2, 2026Ihor Stepanov, Aleksandr Smechov, Mykhailo Shtopko +2LLM RoutingZero-Shot Text Classification

  27. CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing

    Sep 1, 2026Huu Huy Nguyen, Chien Van Nguyen, Franck Dernoncourt +4Long-Context Language Model InferenceLLM Inference Acceleration

  28. CPR for LLMs: Critical-Point Routing against Catastrophic Forgetting in Domain Adaptation

    Aug 31, 2026Kwangmin Ki, Yunhun Nam, Jongheon Jeong +1LLM Fine-TuningDomain Adaptation for LLMs

  29. ProgRouter: Online Progress-Guided Orchestration for Multi-Agent LLM Workflows under Quality-Cost Tradeoffs

    Aug 26, 2026Songyuan Li, Ahmed M. Abdelmoniem, Shiqiang WangLLM Inference EfficiencyMulti-Agent Orchestration

  30. Error-Aware Reverse Auction Mechanism for Large Language Model Routing

    Aug 13, 2026Haolong Chen, Zhengyuan Xin, Liang Zhang +2Cost-Aware InferenceLLM Routing

  31. Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation

    Aug 12, 2026Rodrigo Guedes de Souza, Alison R. PanissonLLM EvaluationLanguage Model Generation Evaluation

  32. TradingMoE: Routing the Right Experts in Evolving Markets

    Aug 12, 2026Chang Zhou, Xingtong Yu, Minbin Huang +4Mixture-of-Experts Language ModelsSparse Mixture-of-Experts

  33. Scheduling Mixed RL Rollouts Beyond Prefix Locality

    Aug 11, 2026Zetao Hong, Song Yuan, Yuanhao Ding +4LLM Inference SchedulingKV-Cache Management

  34. RAISE: Diagnosing Acquisition Collapse in Costly LLM Signals

    Aug 11, 2026Ying Yuan, Yu Wang, Yize Cheng +1Cost-Aware InferenceLLM Routing

  35. MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale

    Aug 11, 2026Yuhang Yao, Zeyu Wang, Wanyi Chen +8Continual Learning for LLM AgentsLLM Fine-Tuning

  36. Opportunity Is Not Realizability: Selection-Valid Diagnostics for Multi-LLM Routing

    Aug 8, 2026Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus SwaqeebModel SelectionLLM Evaluation

  37. LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

    Aug 7, 2026Tao Feng, Fangxu Yu, Haozhen Zhang +9LLM EvaluationLLM Routing

  38. MACRO: Markov Chain Routing of Transformer Layers

    Aug 6, 2026Paweł Batorski, Abtin Pourhadi, Akylgali Aitaza +2Markov ModelsLLM Inference Acceleration

  39. RouteGuard: Certifying Routing Gain in LLM Multi-Agent Systems When Complementarity Is Not Enough

    Aug 5, 2026Anchen Sun, Kaiqi YangMulti-Agent LLM SystemsAI Agent Benchmarks

  40. COMPAS: Difficulty-Aware Joint Search for Optimizing Code Generation

    Aug 5, 2026Jingzhi Gong, Jie M. Zhang, Gunel Jahangirova +3Large Language Model-Guided OptimizationCode Generation

  41. EntropyMoE: Entropy-Aware Sparse Expert Routing for Tokenizer-Free LLMs

    Jul 31, 2026Bo Liu, Muxuab Yu, Yu Zhang +2Sparse Mixture-of-ExpertsByte-Level Language Model