cs.LGSep 30, 2026

Inference Auctions

Authors: Keegan Harris, Siddharth Prasad, Asher Trockman, Nika Haghtalab, Michael I. Jordan

Organizations: University of California, Berkeley · Toyota Technological Institute at Chicago · Google Research · Inria & École Normale Supérieure

Abstract

When inference demand exceeds available compute capacity, model providers must decide which requests should be served first. Users have different tolerances for delay from an LLM API, but current priority pricing schemes compress these differences into coarse fixed-price service tiers. We design an inference auction that allows users to bid for faster service. Our auction allocates priority in an economically efficient way without sacrificing latency, and we develop fast algorithms for implementing prices that incentivize truthful bidding. We also design an autobidding agent for our inference auction, where users specify an inference budget and the autobidder dynamically adjusts its bids over time to maximize user utility subject to the budget constraint. Experiments validate the practicality of our auction: it increases system welfare while maintaining the cache utilization and latency advantages of SGLang, a state-of-the-art inference serving framework.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Learning the Cost of Reliable Inference

    Sep 23, 2026Dimitrios Rontogiannis, Ander Artola Velasco, Manuel Gomez RodriguezQuestion-Answering Benchmarks

  2. The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs

    Jun 2, 2026Xu Wan, Speed Zhu, Jianwei Cai +4Token Budget AllocationEconomies

  3. End-to-End Dynamic Sparsity for Resource-Adaptive LLM Inference

    Jun 26, 2026Yuhang Chen, Jinhao Duan, Ruichen Zhang +11LLM Inference OptimizationSparsity