math.OCAug 25, 2026

The Shadow Price of Intelligence: Quality Degradation in LLM Inference as a Supply Chain Problem

Authors: Elioth Sanabria

Organizations: Department of Decision and Technology Analytics Lehigh University - College of Business

Abstract

Large language model providers are compute constrained, and a common response to congestion is to degrade service. A degraded answer fails with some probability, and a failed answer either returns as a retry or departs as churn, destroying LTV on an unaccounted ledger. We model inference allocation as a newsvendor whose stockout cost is churned lifetime value, a geometric retry multiplier in which the recycled product is dissatisfaction, and a two-regime transient fluid queue whose arrival rate is made endogenous by retries. Statically, there is a regime in which a cheaper model saves energy per initiated task while consuming strictly more capacity per initiated task, so the discount inverts when capacity binds. Dynamically, a reactive throttle fired during a surge can cross an ignition threshold beyond which it manufactures more traffic than it sheds, and a release rule set below the degraded equilibrium converts a transient surge into a permanent degraded regime. With heterogeneous customers, throttling is a transportation problem in retry-inflated load whose optimal policy rations intelligence by critical ratio, and whose dual, the shadow price of intelligence, prices a marginal query by class and by hour. Under congestion, throttling is not a cost lever but a demand lever.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs

    Jun 2, 2026Xu Wan, Speed Zhu, Jianwei Cai +4Token Budget AllocationEconomies

  2. Service-Induced Congestion in Memory-Constrained LLM Serving

    Jun 14, 2026Ruicheng Ao, Jing Dong, Gan Luo +1Large Language Model ServingLarge Language Model Memory

  3. Learning the Cost of Reliable Inference

    Sep 23, 2026Dimitrios Rontogiannis, Ander Artola Velasco, Manuel Gomez RodriguezQuestion-Answering Benchmarks