LLM Routing

LLM: Large Language Model

Momentum

30 papers in the last four weeks, up 88% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 224

All topics
CardsList
  1. Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System

    Oct 6, 2026Panagiotis Kasnesis, Christos Chatzigeorgiou, Lazaros Toumanidis +1Small Language ModelsLLM Agent Evaluation

  2. Do I Need the Cloud? Uncertainty-Aware Step-Level Handoff for Small Language Model Agents

    Oct 6, 2026Abolfazl YounesiSmall Language ModelsLLM Agents

  3. Cross-lingual Calibration of Pre-Generation Success Probes for Multilingual LLM Routing

    Oct 5, 2026Andrea Paganelli, Stefano Civelli, Pietro Bernardelle +1Confidence Estimation in Language ModelsLanguage Model Calibration

  4. Breaking the Tie: A Cluster-Aware Routing Framework for Large Language Models

    Oct 5, 2026Yao Lu, Zhaiyuan Ji, Yaxin Gao +6LLM Routing

  5. OR for AI That Does OR: Routing LLMs up the Escalator inside the OSCAR Framework

    Oct 1, 2026Jinzhi Bu, Haixin Tang, Huanan ZhangLarge Language Model-Guided OptimizationOptimal Stopping

  6. Adaptive-GEPA: Make Your Harness Fit Heterogeneous Requests

    Sep 30, 2026Tianyu Chen, Yasi Zhang, Ruiyi Wang +3Large Language Model-Guided OptimizationLLM-Based Program Synthesis

  7. FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing

    Sep 29, 2026Wang Wei, Harry Yang, Tiankai Yang +6Adaptive Model RoutingLLM Routing

  8. You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference

    Sep 29, 2026Liang He, Jingbo Wen, Yixiong Chen +4LLM ServingCost-Aware Inference

  9. Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models

    Sep 29, 2026Yury Nahshan, Nati Daniel, Jacob Goldberger +1Mixture-of-Experts Language ModelsParameter-Free Mixture-of-Experts Routing

  10. Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing

    Sep 29, 2026Guannan Lai, Gelin Bian, Hao-Xuan Ma +5LLM ServingCost-Aware Inference

  11. Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing

    Sep 29, 2026Guannan Lai, Han-Jia YeLLM Routing

  12. SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing

    Sep 28, 2026Vasilis Perifanis, Nikolaos Pavlidis, Symeon SymeonidisLLM Routing

  13. RSI-Router: Evolving Subtask-Level LLM Routing and Skills for Cost-Efficient Agents

    Sep 28, 2026Hao Li, Hangfan Zhang, Zhiyao Cui +5Cost-Aware InferenceLLM Agent Workflow Optimization

  14. A Persistent State for Auditable Mixture-of-Experts Routing

    Sep 28, 2026Abdurrahman Javat, Allan KazakovLLM AuditingState Tracking

  15. FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents

    Sep 28, 2026Xi Xiao, Yunbei Zhang, Chen Liu +7LLM GroundingLLM Inference Scheduling

  16. MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference

    Sep 28, 2026Junfeng Wu, Zehao Fan, Hadjer Benmeziane +3Expert OffloadingMixture-of-Experts Language Models

  17. Learning the Cost of Reliable Inference

    Sep 23, 2026Dimitrios Rontogiannis, Ander Artola Velasco, Manuel Gomez RodriguezCost-Aware InferenceLLM Routing

  18. PentestChain: A Cost-Aware, MCP-Orchestrated Framework for Automated Penetration Testing with Free-Tier LLMs

    Sep 16, 2026Rushabh Vipulkumar Patel, Dipo Dunsin, Mohammed Almaiah +1Automated Penetration TestingLLMs for Cybersecurity

  19. Beyond the Previous Layer: Residual Predictive Structure in Sparse MoE Routing

    Sep 15, 2026Hao Li, Yasuyuki Tahara, Yuichi SeiSparse Mixture-of-ExpertsExpert Routing

  20. One Size Does Not Fit All! Dynamic Retriever and Generator Selection for RAG

    Sep 15, 2026Neeraj Anand, Payel Santra, Partha Basuchowdhuri +2Retrieval-Augmented GenerationAdaptive RAG

  21. Quality-Constrained Routing over a Fixed Pool of Quantized Mixture-of-Experts Instances

    Sep 14, 2026Zhenghong Huang, Hongfan Wu, Jiheng ZhangMixture-of-Experts InferenceMixture-of-Experts Quantization