cs.AISep 29, 2026

You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference

Authors: Liang He, Jingbo Wen, Yixiong Chen, Yue Yang, Qizhen Lan, Kangning Cui, Xilu Wang

Organizations: University of Sydney · Johns Hopkins University · Stanford University · City University of Hong Kong

Abstract

Existing LLM routers choose among models using static per-model costs. We show that open-weight inference markets introduce a second, largely ignored decision axis: after choosing a model, a client must still choose which provider serves it. Measuring live endpoints across [nummodels] open models, competing providers, multiple task types, and three measurement waves, we find that provider choice cannot be inferred from the price list. The same model can vary sharply in quality, latency, availability, and price across providers; higher-priced providers are consistently faster, but price does not reliably predict quality or availability; and provider feasibility is task-selective, with one deployment nearly normal on knowledge tasks but catastrophically degraded on multi-step reasoning. We formulate same-model provider selection as a price-taker market-aware routing problem. A simple measured-map policy routes to the cheapest provider that is both quality-equivalent and healthy, yielding matched-quality savings while avoiding degraded endpoints. Because the map drifts, we introduce FACET, an online provider router that certifies per-(provider x task) feasibility facets and fails safe to an anchor before serving uncertified arms. Across relaxed deployment assumptions, FACET tolerates imperfect task assignment and sparse feedback, while systematic evaluator bias exposes a quality-signal trust boundary that can be mitigated with ground-truth probes or audits. Live provider runs further confirm that certification can move real traffic from a premium anchor to a substantially cheaper certified endpoint. Our results suggest that market-aware LLM routing must measure not only which model to use, but also who serves it.

Figures & tables

Appendix figures & tables23 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Learning the Cost of Reliable Inference

    Sep 23, 2026Dimitrios Rontogiannis, Ander Artola Velasco, Manuel Gomez RodriguezQuestion-Answering Benchmarks

  2. LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

    Aug 7, 2026Tao Feng, Fangxu Yu, Haozhen Zhang +9Large Language Model RoutingInference Cost

  3. RouteJudge: An Open Platform for Reproducible and Preference-Aware LLM Routing

    Jun 17, 2026Guannan Lai, Haoran Hu, Han-Jia YeLarge Language Model Routing