You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference
Organizations: University of Sydney · Johns Hopkins University · Stanford University · City University of Hong Kong
Abstract
Existing LLM routers choose among models using static per-model costs. We show that open-weight inference markets introduce a second, largely ignored decision axis: after choosing a model, a client must still choose which provider serves it. Measuring live endpoints across [nummodels] open models, competing providers, multiple task types, and three measurement waves, we find that provider choice cannot be inferred from the price list. The same model can vary sharply in quality, latency, availability, and price across providers; higher-priced providers are consistently faster, but price does not reliably predict quality or availability; and provider feasibility is task-selective, with one deployment nearly normal on knowledge tasks but catastrophically degraded on multi-step reasoning. We formulate same-model provider selection as a price-taker market-aware routing problem. A simple measured-map policy routes to the cheapest provider that is both quality-equivalent and healthy, yielding matched-quality savings while avoiding degraded endpoints. Because the map drifts, we introduce FACET, an online provider router that certifies per-(provider x task) feasibility facets and fails safe to an anchor before serving uncertified arms. Across relaxed deployment assumptions, FACET tolerates imperfect task assignment and sparse feedback, while systematic evaluator bias exposes a quality-signal trust boundary that can be mitigated with ground-truth probes or audits. Live provider runs further confirm that certification can move real traffic from a premium anchor to a substantially cheaper certified endpoint. Our results suggest that market-aware LLM routing must measure not only which model to use, but also who serves it.
Figures & tables
| Price vs. | Median Spearman | Interpretation |
| Accuracy | no stable price–quality relation | |
| Latency | higher-priced providers are faster | |
| Availability | no stable price–availability relation |
| Catastrophic exposure | |||
| Scenario | Facet (full) | No anchor | Context-blind |
| S1: Cheap mine from start | 8.3% | 6.7% | |
| S2: Certified arm slips mid-run | 5.6% | 3.2% | |
| S3: No cheap catastrophic mine | 1.2% | 0.0% | — |
| Periodic re-measurement | facet | ||||||
| Probe cost ($) | Below-floor | Total cost ($) | Probe cost ($) | Below-floor | Total cost ($) | ||
| 100 | 1740.2 | 3.70% | 2361 | 0.10 | 16.7 | 1255 | |
| 400 | 487.2 | 0.25 | 22.9 | ||||
| 1600 | 139.2 | 38.60% | 709 | 1.00 | 27.6 | 1130 | |
| Label corruption (known tasks) | OOD traffic | ||||
| Perturbation | Facet | Reference | Perturbation | Facet | Reference |
| error | (context-blind) | OOD | (nearest facet) | ||
| error | OOD | ||||
| error | OOD | ||||
| error | |||||
| Metric | Value | Metric | Value |
| Deployment scale | |||
| Served queries | Certification probes | ||
| Failed calls | Aggregate served accuracy | ||
| Serving behavior | |||
| Avg. price, queries 1–100 | \mathbf{\0.924/M}$ | Avg. price, queries 301–500 | \mathbf{\0.210/M}$ |
| Anchor price | $1.04 / M | ||
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Decision axis | Cost model | Cross-provider quality | |
| FrugalGPT / RouteLLM | |||
| CARROT / MixLLM | which model | static list price | — |
| PriLLM (provider-side) | sets its own price | endogenous | — |
| Token Arena (measurement) | — (leaderboard) | — | static snapshot |
| Ours | which provider | live, exogenous | measured, live |
| Cost premium | Below-floor rate | Quality-regret | |||
| Policy | vs. oracle | Median | Worst | Median | Worst |
| Static-argmin | |||||
| Cost-Thompson | |||||
| SW-UCB | |||||
| M-UCB (change-detect) | |||||
| facet core | |||||
| Policy outcome | Break-even threshold | |||
| Baseline | Cost ($) | Below-floor serves | ($) | Query-price equiv. |
| Frozen map | ||||
| Periodic | ||||
| M-UCB | ||||
| SW-UCB | ||||
| Latency | Accuracy | Availability | ||||
| Run selection | Median | Negative cells | Median | Negative cells | Median | Negative cells |
| Latest | ||||||
| Earliest | ||||||
| Across-run mean | ||||||
| Latency component | Median with price | Negative cells |
| Seconds per generated token | ||
| Fixed request overhead |
| Statistic | Value |
| Evaluation set | |
| Cells with at least two feasible providers | |
| Latency consequence | |
| Median normalized speed rank of cheapest feasible provider | |
| Cheapest feasible provider is strictly slowest | |
| Median p95-latency penalty vs. fastest feasible provider | |
| Routing outcome | |||
| Latency deadline | Feasible cells | Relative cost | Route changes |
| s | |||
| s | |||
| s | |||
| s | |||
| s | |||
| Safety outcome | Monitoring cost | ||
| Probe rate | Below-floor median | IQR | Probe spend median ($) |
| [ , ]% | |||
| [ , ]% | |||
| [ , ]% | |||
| [ , ]% | |||
| [ , ]% | |||
| Measured-map router | OOS baselines | ||||
| Health threshold | Cells | In-sample | OOS | Cheapest | Premium |
| Quality margin: 2 points | |||||
| Quality margin: 5 points | |||||
| facet core | Best baseline | |||
| Drift regime | Below-floor | q-regret | Below-floor | q-regret |
| Abrupt-change regimes | ||||
| Abrupt, spaced (assumed) | ||||
| Frequent (sub-recovery) | ||||
| Correlated cohort slip | ||||
| Outside the abrupt-change assumption | ||||
| Policy | Mean acc. | Rel. cost ∗ | Below-floor | Mean avail. |
| Single-best (quality-first) | ||||
| Premium (most expensive) | ||||
| Random | ||||
| Cheapest (price-blind) | ||||
| Ours |
| In-sample | Out-of-sample | |||||
| Policy | Accuracy | Below-floor | q-regret | Accuracy | Below-floor | q-regret |
| Single-best | ||||||
| Router (ours) | ||||||
| Cheapest | ||||||
| Premium | ||||||
| Cell subset | Cells | Median saving |
| All cells | ||
| Saturated cells only | ||
| Non-saturated cells only |
| Model | Task | Saving | Providers |
| llama-3.1-8b-instruct | extraction | ||
| math | |||
| llama-3.3-70b-instruct | extraction | ||
| math | |||
| llama-4-maverick | extraction | ||
| math |
| Cost outcome | Safety outcome | |||
| Policy | Cost ($) | Below-floor | Catastrophic | q-regret / |
| S1: Cheap mine from the start | ||||
| Periodic | ||||
| Periodic | ||||
| Periodic | ||||
| SW-UCB | ||||
| Cell | Route before | Route after | Driver |
| Interval 1 (days 0–13) | |||
| deepseek-v3.1 / extraction | AtlasCloud | DeepInfra | availability recovery |
| deepseek-v3.1 / math | AtlasCloud | DeepInfra | availability recovery |
| llama-3.3-70B / extraction | DeepInfra | Nebius | quality slip below floor |
| mistral-small / classification | Venice | DeepInfra | rate limiting |
| Interval 2 (days 13–43) | |||
| Catastrophic exposure | ||
| Task-label error | facet | Context-blind |
| Catastrophic exposure | ||
| OOD share | Unknown fail-safe | Nearest known facet |
| Safety outcome | Cost outcome | |
| Judge-label error | Below-floor | Total cost ($) |
| Below-floor exposure | ||
| Observed-feedback fraction | facet | Reference |
| — | ||
| — | ||
| — | ||
| SW-UCB: | ||
| Below-floor exposure | ||||
| Setting | facet | Periodic | SW-UCB | Static |
| Clean | — | |||
| noise labels | ||||
| noise labels probe rate | ||||
| Cost outcome | Quality and safety outcome | |||
| Provider policy | Cost ($) | Accuracy | Below-floor | Catastrophic |
| GSM8K | ||||
| Dearest | ||||
| Cheapest | ||||
| Ours | ||||
| MMLU | ||||
| Market dynamics | ||||
| Model | Providers | Price changes | Re-elections | Re-elections/day |
| deepseek-chat-v3.1 | ||||
| llama-3.3-70b-instruct | ||||
| deepseek-v4-pro-0813 | ||||
| deepseek-v4-flash-0731 | ||||