Cost-Aware Inference

Latest papers 183

All topics
CardsList
  1. AutoRelAnnotator: Calibrated Model Cascades for Cost-Efficient Relevance Evaluation in Sponsored Search

    Jun 24, 2026Md Omar Faruk Rokon, Shasvat Desai, Hong Yao +1Online AdvertisingCost-Aware Inference

  2. Cost-Optimal Decision Diagrams for Stochastic Boolean Function Evaluation

    Jun 23, 2026Xia Zong, Tuomo Lehtonen, Jussi RintanenBranch and BoundCost-Aware Inference

  3. Token-Operations-Oriented Inference Optimization Techniques for Large Models

    Jun 18, 2026Shiguo Lian, Kai Wang, Zhaoxiang Liu +22LLM ServingInference-Time Optimization

  4. Think Again or Think Longer? Selective Verification for Budget-Aware Reasoning

    Jun 18, 2026Sajib Acharjee Dip, Dawei Zhou, Liqing ZhangLLM Inference EfficiencyCost-Aware Inference

  5. Closing the Operational Gap in Semantic Caching

    Jun 18, 2026Aditeya Baral, Radoslav Ralev, Iliya Sotirov Zhechev +2LLM EvaluationCost-Aware Inference

  6. RouteJudge: An Open Platform for Reproducible and Preference-Aware LLM Routing

    Jun 17, 2026Guannan Lai, Haoran Hu, Han-Jia YeLLM EvaluationPairwise Preference Evaluation

  7. Online LLM Selection via Constrained Bandits with Time-Varying Demand

    Jun 16, 2026Yin Huang, Qingsong Liu, Jie XuMulti-Armed BanditsCost-Aware Inference

  8. ConCise: Training-Free Conclusion-Chain State Compression for Cost-Efficient Multi-Step RAG Services

    Jun 13, 2026Kuan Yan, Zhiqing Tang, Tian Wang +1Retrieval-Augmented GenerationCost-Aware Inference

  9. Are Online Skill and Memory Modules Always Worth Their Tokens? A Budget-Constrained Study of Web Agents

    Jun 12, 2026Sina Hajimiri, Masih Aminbeidokhti, Jose Dolz +4Web Agent BenchmarksLLM Agent Evaluation

  10. PLAIground: SLO-Driven Runtime Model Selection for Compound AI Systems in the Edge-Cloud-Space Continuum

    Jun 12, 2026Milos Gravara, Cynthia Marcelino, Andrija Stanisic +1Cost-Aware InferenceEdge-Cloud Computing

  11. Design Methodology and Performance Trade-offs Management for Distributed and Compound AI Systems

    Jun 12, 2026Milos Gravara, Andrija Stanisic, Stefan NasticCost-Aware Inference

  12. Cost-Optimal LLM Routing with Limited User Feedback under User Satisfaction Guarantees

    Jun 12, 2026Herbert Woisetschläger, Arastun Mammadli, Ryan Zhang +1Cost-Aware InferenceLLM Routing

  13. Can I Buy Your KV Cache?

    Jun 11, 2026Luoyuan ZhangLLM Inference EfficiencyKV Caching

  14. AI Tokenomics: The Economics of Tokens, Computation, and Pricing in Foundation Models

    Jun 10, 2026Quanyan ZhuCost-Aware InferenceToken Efficiency

  15. Belief-Space Control for Personalized Cancer Treatment via Active Inference

    Jun 9, 2026Deniz Sargun, H. Bugra Tulay, C. Emre KoksalBelief-Space PlanningHealthcare

  16. BUDDY: BUdget-Driven DYnamic Depth Routing for Adaptive Large Language Model Inference

    Jun 8, 2026Yuhua Zhou, Shaoqi Yu, Shichao Weng +4LLM PruningCost-Aware Inference

  17. Bayesian Selective Latent Inference for Wastewater-First Influenza Monitoring

    Jun 8, 2026Yixuan Zhang, Yang Song, Hao Wang +2Epidemic ForecastingSelective Prediction

  18. FAME: Forecastability-Aware Mixture of Experts for Heterogeneous Time Series Forecasting

    Jun 8, 2026Qianyang Li, Xingjun Zhang, Shaoxun Wang +2Sparse Mixture-of-ExpertsTime Series Forecasting

  19. Larch: Learned Query Optimization for Semantic Predicates

    Jun 6, 2026Fuheng Zhao, Pawel Liskowski, Zihan Li +5Cost-Aware Inference

  20. Learning to Route LLMs from Implicit Cost-Performance Preferences via Meta-Learning

    Jun 4, 2026Jiahao Zeng, Ming Tang, Ningning DingMeta-LearningCost-Aware Inference

  21. CRAFT: Cost-aware Refinement And Front-aware Tuning of Prompts

    Jun 3, 2026Shanu Kumar, Shubhanshu Khandelwal, Akhila Yesantarao Venkata +3LLM PromptingCost-Aware Inference

  22. Not All Errors Are Equal: Consequence-Aware Reasoning Compute Allocation

    Jun 3, 2026Liang He, Jingbo Wen, Haoyu Wang +4Cost-Aware InferenceInference-Time Compute Allocation

  23. NetKV: Network-Aware Decode Instance Selection for Disaggregated LLM Inference

    Jun 2, 2026Mubarak Adetunji OjewaleLLM InferenceKV Caching

  24. Resource-Constrained Adaptive Inference for Sequential Pricing

    Jun 2, 2026Ruicheng Ao, Jiashuo Jiang, David Simchi-LeviAdaptive InferenceCost-Aware Inference

  25. The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs

    Jun 2, 2026Xu Wan, Speed Zhu, Jianwei Cai +4Cost-Aware InferenceInference-Time Scaling