Existing LLM routers choose among models using static per-model costs. We show that open-weight inference markets introduce a second, largely ignored decision axis: after choosing a model, a client must still choose which provider serves it. Measuring live endpoints across [nummodels] open models, competing providers, multiple task types, and three measurement waves, we find that provider choice cannot be inferred from the price list. The same model can vary sharply in quality, latency, availability, and price across providers; higher-priced providers are consistently faster, but price does not reliably predict quality or availability; and provider feasibility is task-selective, with one deployment nearly normal on knowledge tasks but catastrophically degraded on multi-step reasoning. We formulate same-model provider selection as a price-taker market-aware routing problem. A simple measured-map policy routes to the cheapest provider that is both quality-equivalent and healthy, yielding matched-quality savings while avoiding degraded endpoints. Because the map drifts, we introduce FACET, an online provider router that certifies per-(provider x task) feasibility facets and fails safe to an anchor before serving uncertified arms. Across relaxed deployment assumptions, FACET tolerates imperfect task assignment and sparse feedback, while systematic evaluator bias exposes a quality-signal trust boundary that can be mitigated with ground-truth probes or audits. Live provider runs further confirm that certification can move real traffic from a premium anchor to a substantially cheaper certified endpoint. Our results suggest that market-aware LLM routing must measure not only which model to use, but also who serves it.
Figures & tables
Price vs.
Median Spearman
Interpretation
Accuracy
+0.05
no stable price–quality relation
Latency
−0.61
higher-priced providers are faster
Availability
0.00
no stable price–availability relation
Table 2: Benchmark-scale within-(model × task) Spearman correlations between provider price and metrics. Price is strongly associated with lower latency, but not with accuracy or availability.
Catastrophic exposure
Scenario
Facet (full)
No anchor
Context-blind
S1: Cheap mine from start
0.0%
8.3%
6.7%
S2: Certified arm slips mid-run
0.4%
5.6%
3.2%
S3: No cheap catastrophic mine
1.2%
0.0%
—
Table 3: Ablation experiments isolating fail-safe anchoring and task-conditioned certification. Values report catastrophic exposure (below-floor served queries).
Periodic re-measurement
facet
P
Probe cost ($)
Below-floor
Total cost ($)
r
Probe cost ($)
Below-floor
Total cost ($)
100
1740.2
3.70%
2361
0.10
16.7
0.17%
1255
400
487.2
14.30%
1089
0.25
22.9
0.67%
1110
1600
139.2
38.60%
709
1.00
27.6
0.28%
1130
Table 4: Periodic re-measurement versus facet . A certified provider can slip between measurement rounds. Periodic re-probing reduces exposure by probing more frequently, while facet continuously monitors served facets and uses probes only for certification (10 seeds).
Label corruption (known tasks)
OOD traffic
Perturbation
Facet
Reference
Perturbation
Facet
Reference
0% error
0.00%
6.72% (context-blind)
10% OOD
0.00%
3.41% (nearest facet)
10% error
2.23%
6.72%
25% OOD
0.00%
1.03%
20% error
3.54%
6.72%
50% OOD
0.00%
0.24%
33% error
4.72%
6.72%
Table 5: Robustness to imperfect task assignment. Left: corruption of known task labels. Right: OOD traffic treated as uncertified or mapped to the nearest known facet. Values report catastrophic exposure (below-floor served queries; 10 seeds).
Metric
Value
Metric
Value
Deployment scale
Served queries
538
Certification probes
130
Failed calls
4
Aggregate served accuracy
88.3%
Serving behavior
Avg. price, queries 1–100
\mathbf{\0.924/M}$
Avg. price, queries 301–500
\mathbf{\0.210/M}$
Anchor price
$1.04 / M
Table 6: Live facet deployment: 36-hour run on Llama-3.3-70B across 12 providers. Validates cold-start certification and migration; no quality slip occurred during the window.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Decision axis
Cost model
Cross-provider quality
FrugalGPT / RouteLLM
CARROT / MixLLM
which model
static list price
—
PriLLM (provider-side)
sets its own price
endogenous
—
Token Arena (measurement)
— (leaderboard)
—
static snapshot
Ours
which provider
live, exogenous
measured, live
Appendix
Table 7: Positioning against representative routing and measurement work. Prior client-side routers select among models; our layer selects among providers serving the same selected open-weight model.
Cost premium
Below-floor rate
Quality-regret
Policy
vs. oracle
Median
Worst
Median
Worst
Static-argmin
−6%
33.0%
33.0%
49.50
49.50
Cost-Thompson
+4%
6.5%
47.6%
9.08
20.45
SW-UCB
+13%
2.4%
56.6%
3.33
6.26
M-UCB (change-detect)
+14%
2.5%
53.2%
3.71
6.17
facet core
+59%
1.3%
42.6%
1.78
4.41
Appendix
Table 8: facet ’s detector core versus non-stationary bandit baselines. Results cover 12 benchmark cells against a clairvoyant cheapest-feasible oracle. Each safety metric is shown at the median and at each policy’s own worst cell. Cost premium is relative to the oracle; quality-regret is the summed below-floor shortfall ∑tmax(0,θ−qat) per 103 serves.
Policy outcome
Break-even threshold
Baseline
Cost ($)
Below-floor serves
λ⋆ ($)
Query-price equiv.
Frozen map
1015
990
0.78
1.5×
Periodic P=400
1448
595
0.40
0.8×
M-UCB
1241
314
6.02
11×
SW-UCB
1230
296
9.08
17×
Appendix
Table 9: Break-even cost of avoiding one below-floor serve. facet is preferred whenever the deployment-level loss associated with one below-floor answer exceeds λ⋆ . The final column expresses the threshold in units of the mean price of an ordinary served query.
Latency
Accuracy
Availability
Run selection
Median ρ
Negative cells
Median ρ
Negative cells
Median ρ
Negative cells
Latest
−0.610
21/21
+0.049
10/21
+0.000
9/21
Earliest
−0.610
21/23
−0.224
14/23
+0.000
10/23
Across-run mean
−0.566
22/23
−0.063
13/23
+0.000
11/23
Appendix
Table 11: Robustness of within-cell price correlations to measurement-run selection. The correlation direction and the number of cells with negative correlation are reported for each service metric.
Latency component
Median ρ with price
Negative cells
Seconds per generated token
−0.675
21/21
Fixed request overhead
+0.274
7/21
Appendix
Table 12: Decomposition of the price–latency relationship. Correlations are computed across providers within each model–task cell.
Statistic
Value
Evaluation set
Cells with at least two feasible providers
17
Latency consequence
Median normalized speed rank of cheapest feasible provider
1.00
Cheapest feasible provider is strictly slowest
9/17
Median p95-latency penalty vs. fastest feasible provider
7.8×
Appendix
Table 13: Latency consequence of choosing the cheapest quality-feasible provider.
Routing outcome
Latency deadline
Feasible cells
Relative cost
Route changes
30 s
12/12
1.00×
0%
10 s
12/12
1.19×
58%
5 s
12/12
1.31×
58%
2 s
10/12
2.03×
70%
1 s
6/12
2.62×
100%
Appendix
Table 14: Latency-constrained measured-map routing. Relative cost is normalized to the unconstrained cheapest quality-feasible route. Route changes report the fraction of feasible cells whose selected provider changes after imposing the deadline.
Safety outcome
Monitoring cost
Probe rate
Below-floor median
IQR
Probe spend median ($)
0.10
0.00%
[ 0.00 , 0.80 ]%
17.0
0.25
0.44%
[ 0.00 , 0.80 ]%
17.6
0.50
0.60%
[ 0.00 , 0.88 ]%
28.2
1.00
0.14%
[ 0.00 , 0.96 ]%
29.8
2.00
0.92%
[ 0.52 , 1.36 ]%
48.9
Appendix
Table 15: Run-to-run variability of the facet probe-rate sweep. Results are computed over 20 seeds. The overlapping interquartile ranges at low probe rates show that the ordering of nearby point estimates should not be interpreted as a tuning effect.
Measured-map router
OOS baselines
Health threshold
Cells
In-sample
OOS
Cheapest
Premium
Quality margin: 2 points
80%
9
0.0%
22.2%
44.4%
55.6%
90%
9
0.0%
22.2%
44.4%
55.6%
Quality margin: 5 points
80%
9
0.0%
11.1%
22.2%
33.3%
Appendix
Table 16: Out-of-sample sensitivity to the quality margin and health threshold. Results cover nine benchmark cells. The earlier wave determines the route and the later wave evaluates it. Values are below-floor rates.
facet core
Best baseline
Drift regime
Below-floor
q-regret
Below-floor
q-regret
Abrupt-change regimes
Abrupt, spaced (assumed)
1.3%
1.9
1.4%
2.9
Frequent (sub-recovery)
1.3%
2.0
2.5%
4.5
Correlated cohort slip
3.3%
3.9
16.4%
21.0
Outside the abrupt-change assumption
Appendix
Table 17: facet ’s detector core across drift shapes. Values report below-floor rate and quality-regret on the flagship cell. The detector is robust to frequent and correlated drift, but the advantage does not hold under gradual ramps or multiple concurrent mines, which are outside its abrupt-change assumption.
Policy
Mean acc.
Rel. cost ∗
Below-floor
Mean avail.
Single-best (quality-first)
95%
1.66×
0%
96%
Premium (most expensive)
92%
3.67×
17%
95%
Random
92%
2.19×
22%
98%
Cheapest (price-blind)
91%
1.00×
28%
97%
Ours
94%
1.57×
0%
99%
Appendix
Table 18: Routing policy comparison over all measured model–task cells. Relative cost normalizes each cell to its cheapest provider. Below-floor is the fraction of cells whose selected provider falls below the cell’s quality floor.
In-sample
Out-of-sample
Policy
Accuracy
Below-floor
q-regret
Accuracy
Below-floor
q-regret
Single-best
87%
0%
0.0
87%
0%
0.0
Router (ours)
85%
0%
0.0
86%
11%
0.0
Cheapest
84%
22%
0.1
85%
22%
0.1
Premium
79%
44%
4.3
80%
33%
4.0
Appendix
Table 19: Out-of-sample routing. Providers are selected using an earlier measurement wave and evaluated on a later wave with n≥100 benchmark accuracy.
Cell subset
Cells
Median saving
All cells
18
56.5%
Saturated cells only
9
56.5%
Non-saturated cells only
9
50.0%
Appendix
Table 20: Matched-quality savings after separating saturated and non-saturated model–task cells.
Model
Task
Saving
Providers
llama-3.1-8b-instruct
extraction
84.1%
5
math
84.1%
5
llama-3.3-70b-instruct
extraction
79.8%
10
math
79.2%
11
llama-4-maverick
extraction
50.0%
4
math
50.0%
4
Appendix
Table 21: Complete per-cell results for the nine non-saturated cells. Saving is the reduction of the selected measured-equivalent healthy route relative to the premium provider.
Cost outcome
Safety outcome
Policy
Cost ($)
Below-floor
Catastrophic
q-regret / 103
S1: Cheap mine from the start
Periodic P=200
1834.8
2.78%
2.78%
8.33
Periodic P=500
1212.8
0.00%
0.00%
0.00
Periodic P=1000
1004.2
0.00%
0.00%
0.00
SW-UCB
1040.3
1.79%
1.79%
5.52
Appendix
Table 22: Periodic re-probing under three drift regimes. S1 contains a cheap mine from the start; S2 lets a certified provider slip mid-run; S3 contains repeated measured-style changes.
Cell
Route before
Route after
Driver
Interval 1 (days 0–13)
deepseek-v3.1 / extraction
AtlasCloud
DeepInfra
availability recovery
deepseek-v3.1 / math
AtlasCloud
DeepInfra
availability recovery
llama-3.3-70B / extraction
DeepInfra
Nebius
quality slip below floor
mistral-small / classification
Venice
DeepInfra
rate limiting
Interval 2 (days 13–43)
Appendix
Table 25: Decision changes across three measurement waves. Despite mostly stable prices, the cheapest feasible provider changes because of health, availability, provider churn, and quality drift.
Catastrophic exposure
Task-label error
facet
Context-blind
0%
0.00%
6.72%
5%
1.06%
6.72%
10%
2.23%
6.72%
20%
3.54%
6.72%
33%
4.72%
6.72%
Appendix
Table 26: Task-label corruption. Exposure rises gradually as the task label becomes less reliable.
Catastrophic exposure
OOD share
Unknown → fail-safe
Nearest known facet
10%
0.00%
3.41%
25%
0.00%
1.03%
50%
0.00%
0.24%
Appendix
Table 27: Unknown-task traffic. Fail-safe routing avoids borrowing certification from an unrelated known task.
Safety outcome
Cost outcome
Judge-label error
Below-floor
Total cost ($)
0%
0.28%
1130
20%
0.00%
2868
30%
0.00%
2942
Appendix
Table 28: Symmetric quality-label noise. Random label flips primarily raise fallback cost rather than below-floor exposure. Results average over 10 seeds.
Below-floor exposure
Observed-feedback fraction
facet
Reference
100%
0.28%
—
50%
1.26%
—
20%
1.47%
—
5%
4.57%
SW-UCB: 13.32%
Appendix
Table 30: Sparse quality feedback. Safety degrades as labeled traffic becomes rare, but remains substantially better than the comparison policies at the lowest label rate. Results average over 10 seeds.
Below-floor exposure
Setting
facet
Periodic
SW-UCB
Static
Clean
0.28%
14.30%
2.16%
—
20% noise +20% labels
0.14%
17.00%
10.52%
65.00%
20% noise +20% labels + probe rate 0.25
0.00%
17.00%
10.52%
65.00%
Appendix
Table 31: Joint information stress test. The combination of noisy and sparse feedback sharply reduces cost efficiency, but does not produce a comparable collapse in safety.
Cost outcome
Quality and safety outcome
Provider policy
Cost ($)
Accuracy
Below-floor
Catastrophic
GSM8K
Dearest
1435.7
85.4%
24.6%
24.6%
Cheapest
222.6
92.5%
0.0%
0.0%
Ours
222.6
92.5%
0.0%
0.0%
MMLU
Appendix
Table 32: Provider routing underneath model routing. Cost is reported over 3000 replayed queries. When the cheapest provider is already safe, our layer leaves the route unchanged. When it is degraded, provider-aware selection trades a small cost increase over the cheapest default for substantially lower below-floor exposure.
Market dynamics
Model
Providers
Price changes
Re-elections
Re-elections/day
deepseek-chat-v3.1
8
0
0
0.0
llama-3.3-70b-instruct
12
0
0
0.0
deepseek-v4-pro-0813
15
2
2
2.0
deepseek-v4-flash-0731
30
14
3
3.0
Appendix
Table 33: High-frequency price monitoring over 24 hours at 15-minute resolution. A re-election is a change in the identity of the cheapest provider for that model. Price movement and re-election frequency both increase with the number of competing providers.
School of Artificial Intelligence, Nanjing University, China · 2National Key Laboratory for Novel Software Technology, Nanjing University, 210023, China · 3SinapisAI