Cost-Aware Inference

Latest papers 183

All topics
CardsList
  1. Flexible Routing via Uncertainty Decomposition

    May 8, 2026Charlotte Peale, Siddartha Devic, Parikshit Gopalan +2Uncertainty QuantificationAdaptive Model Routing

  2. PaT: Planning-after-Trial for Efficient Test-Time Code Generation

    May 8, 2026Youngsik Yoon, Sungjae Lee, Seockbean Song +3Test-Time OptimizationExecution-Guided Code Generation

  3. Switchcraft: AI Model Router for Agentic Tool Calling

    May 8, 2026Sharad Agarwal, Pooria Namyar, Alec Wolman +3Tool-Augmented Language Model AgentsCost-Aware Inference

  4. Policy-Guided Stepwise Model Routing for Cost-Effective Reasoning

    May 7, 2026Wenwen Si, Insup Lee, Osbert BastaniCost-Aware InferenceEfficient Language Model Reasoning

  5. On Time, Within Budget: Constraint-Driven Online Resource Allocation for Agentic Workflows

    May 7, 2026Xinglin Wang, Zishen Liu, Shaoxiong Feng +9Stochastic OptimizationCost-Aware Inference

  6. Inference-Time Budget Control for LLM Search Agents

    May 7, 2026Zhengru Fang, Senkang Forest Hu, Zhonghao Chang +6Multi-Hop QAInference-Time Search

  7. Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs

    May 5, 2026Yixuan Mei, Zikun Li, Zixuan Chen +5LLM ServingCost-Aware Inference

  8. Joint Energy Management and Coordinated AIGC Workload Scheduling for Distributed Data Centers: A Diffusion-Aided Reward Shaping Approach

    May 3, 2026Yang Fu, Peng Qin, Liming Chen +3Reward ShapingCost-Aware Inference

  9. Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling

    May 2, 2026Florian Valentin Wunderlich, Lars Benedikt Kaesberg, Jan Philip Wahle +2Self-Consistency DecodingCost-Aware Inference

  10. Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference

    May 1, 2026Yuxuan Gao, Megan Wang, Yi Ling YuLLM EvaluationEnergy-Efficient ML

  11. Rethinking Network Topologies for Cost-Effective Mixture-of-Experts LLM Serving

    Apr 30, 2026Junsun Choi, Sam Son, Sunjin Choi +5LLM ServingCost-Aware Inference

  12. AI Inference as Relocatable Electricity Demand: A Latency-Constrained Energy-Geography Framework

    Apr 30, 2026Xubin Luo, Cheng YangEnergy-Efficient MLCost-Aware Inference

  13. Belief-Guided Inference Control for Large Language Model Services via Verifiable Observations

    Apr 30, 2026Wenhao Yuan, Chenchen Lin, Jian Chen +3LLM InferenceCost-Aware Inference

  14. CGU-ILALab at FoodBench-QA 2026: Comparing Traditional and LLM-based Approaches for Recipe Nutrient Estimation

    Apr 28, 2026Wei-Chun Chen, Yu-Xuan Chen, I-Fang Chung +1LLM EvaluationCost-Aware Inference

  15. RouteNLP: Closed-Loop LLM Routing with Conformal Cascading and Distillation Co-Optimization

    Apr 26, 2026Dongxin Guo, Jikun Wu, Siu Ming YiuLLM ServingCost-Aware Inference

  16. MTRouter: Cost-Aware Multi-Turn LLM Routing with History-Model Joint Embeddings

    Apr 26, 2026Yiqun Zhang, Hao Li, Zihan Wang +6Cost-Aware InferenceLLM Routing

  17. How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks

    Apr 24, 2026Longju Bai, Zhemin Huang, Xingyao Wang +5Coding AgentsCost-Aware Inference

  18. QuantClaw: Precision Where It Matters for OpenClaw

    Apr 24, 2026Manyi Zhang, Ji-Fu Li, Zhongao Sun +5Mixed-Precision QuantizationLLM Quantization

  19. Continuous Semantic Caching for Low-Cost LLM Serving

    Apr 21, 2026Baran Atalar, Xutong Liu, Jinhang Zuo +3LLM ServingCost-Aware Inference

  20. Are Large Language Models Economically Viable for Industry Deployment?

    Apr 21, 2026Abdullah Mohammad, Sushant Kumar Ray, Pushkar Arora +5LLM Inference EfficiencyLLM Evaluation

  21. When Spike Sparsity Does Not Translate to Deployed Cost: VS-WNO on Jetson Orin Nano

    Apr 18, 2026Jason Yoo, Shailesh Garg, Souvik Chakraborty +1Efficient Neural Network InferenceActivation Sparsity

  22. Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization

    Apr 16, 2026Haochun Tang, Yuliang Yan, Jiahua Lu +2Adversarial Prompt GenerationCost-Aware Inference

  23. Streaming Model Cascades for Semantic SQL

    Apr 1, 2026Paweł Liskowski, Kyle SchmausCost-Aware InferenceLLM Routing

  24. More Bang for the Buck: Improving the Inference of Large Language Models at a Fixed Budget using Reset and Discard (ReD)

    Jan 29, 2026Sagi Meir, Tommer D. Keidar, Noam Levi +2LLM Inference EfficiencyLLM Evaluation