LLM Inference Efficiency

LLM: Large Language Model

Latest papers 234

All topics
CardsList
  1. Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation

    May 8, 2026Joon Ha Kim, Geon-Woo Kim, Anoop Rachakonda +1LLM Inference EfficiencyLLM Inference

  2. Learning Agent Routing From Early Experience

    May 8, 2026Yimin Wang, Jiahao Qiu, Xuan Qi +6LLM Inference EfficiencyLLM Agent Routing

  3. Efficient Serving for Dynamic Agent Workflows with Prediction-based KV-Cache Management

    May 7, 2026Haoyu Zheng, Fangcheng Fu, Jia Wu +6LLM Inference EfficiencyKV Caching

  4. One Pool, Two Caches: Adaptive HBM Partitioning for Accelerating Generative Recommender Serving

    May 6, 2026Wenjun Yu, Shuguang Han, Amelie Chi ZhouLLM Inference EfficiencyGPU Acceleration

  5. Tile-Level Activation Overlap for Efficient LLM Inference

    May 5, 2026Abhinav Jangda, Tyler Sorensen, Sebastian Burckhardt +3LLM Inference EfficiencyGPU Kernel Optimization

  6. Human-Less LLM Serving: Quantifying the Human Tax on Throughput

    May 3, 2026Jianhui Lian, Li Chen, Dan Li +1LLM Inference EfficiencyLLM Serving

  7. AgentStop: Terminating Local AI Agents Early to Save Energy in Consumer Devices

    May 1, 2026Dzung Pham, Kleomenis Katevas, Ali Shahin Shamsabadi +1LLM Inference EfficiencyOn-Device Language Model Inference

  8. LLM-Guided Runtime Parameter Optimization for Energy-Efficient Model Inference

    Apr 29, 2026Katelyn Crumpacker, Dimitrios NikolopoulosLLM Inference EfficiencyEnergy-Efficient ML

  9. Evergreen: Efficient Claim Verification for Semantic Aggregates

    Apr 28, 2026Alexander W. Lee, Benjamin Han, Shayak Sen +3LLM Inference EfficiencyClaim Verification

  10. Pythia: Exploiting Workflow Predictability for Efficient Agent-Native LLM Serving

    Apr 28, 2026Shan Yu, Junyi Shu, Yuanjiang Ni +14LLM Inference EfficiencyMulti-Agent LLM Systems

  11. Latency and Cost of Multi-Agent Intelligent Tutoring at Scale

    Apr 27, 2026Iizalaarab Elhaimeur, Nikos ChrisochoidesLLM Inference EfficiencyMulti-Agent LLM Systems

  12. Scalable LLM-based Coding of Dialogue in Healthcare Simulation: Balancing Coding Performance, Processing Time, and Environmental Impact

    Apr 25, 2026Kiyoshige Garces, Gloria Milena Fernandez-Nieto, Linxuan Zhao +4LLM Inference EfficiencyLLM Prompting

  13. PExA: Parallel Exploration Agent for Complex Text-to-SQL

    Apr 24, 2026Tanmay Parekh, Ella Hofmann-Coyle, Shuyi Wang +3LLM Inference EfficiencyText-to-SQL

  14. EverydayGPT: Confidence-Gated Routing for Efficient and Safe Hybrid GPT-RAG Conversational QA

    Apr 24, 2026Jaspreet Singh NahalLLM Inference EfficiencyRetrieval-Augmented Generation

  15. Sub-Token Routing for KV Cache Compression

    Apr 23, 2026Wei Jiang, Wei WangLLM Inference EfficiencyKV Caching

  16. TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping

    Apr 22, 2026Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo +1LLM Inference EfficiencyEarly Stopping

  17. RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory

    Apr 22, 2026Fei Zuo, Zikang Zhou, Hao Cong +2LLM Inference EfficiencyMixed-Precision Quantization

  18. Are Large Language Models Economically Viable for Industry Deployment?

    Apr 21, 2026Abdullah Mohammad, Sushant Kumar Ray, Pushkar Arora +5LLM Inference EfficiencyLLM Evaluation

  19. How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning

    Apr 21, 2026Zhiyuan Zhai, Xinkai You, Wenjing Yan +1LLM Inference EfficiencyRL for Language Model Reasoning

  20. Learning to Seek Help: Dynamic Collaboration Between Small and Large Language Models

    Apr 20, 2026Hang Zeng, Xiangyu Liu, Yong Hu +5LLM Inference EfficiencySmall Language Models

  21. SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving

    Apr 19, 2026Christian LysenstøenLLM Inference EfficiencyLLM Serving

  22. ONTO: A Token-Efficient Columnar Notation for LLM Input Optimization

    Apr 19, 2026Harshavardhanan DeekeswarLLM Inference EfficiencyToken Efficiency

  23. Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning

    Apr 19, 2026Raman Saparkhan, Majd Hawasly, Md Rizwan Parvez +1LLM Inference EfficiencySelf-Consistency Decoding

  24. Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling

    Apr 19, 2026Zizhang Luo, Yuhao Luo, Youwei Xiao +3LLM Inference EfficiencyMulti-Agent LLM Systems

  25. HiveMind: OS-Inspired Scheduling for Concurrent LLM Agent Workloads

    Apr 18, 2026Justice Owusu Agyemang, Jerry John Kponyo, Obed Kwasi Somuah +3LLM Inference EfficiencyAI Coding Agents

  26. When Agents Go Quiet: Output Generation Capacity and Format-Cost Separation for LLM Document Synthesis

    Apr 17, 2026Justice Owusu Agyemang, Michael Agyare, Miriam Kobbinah +2LLM Inference EfficiencyLong-Form Document Generation

  27. PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving

    Apr 14, 2026Xu Bai, Muhammed Tawfiqul Islam, Chen Wang +1LLM Inference EfficiencyLLM Serving

  28. When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models

    Mar 27, 2026Juan Gabriel Kostelec, Qinghai GuoLLM Inference EfficiencyLLM Evaluation

  29. Clinical Note Bloat Reduction for Efficient LLM Use

    Mar 21, 2026Jordan L. Cahoon, Chloe Stanwyck, Asad Aali +6LLM Inference EfficiencyElectronic Health Records

  30. MoEless: Efficient MoE LLM Serving with Serverless Experts

    Mar 6, 2026Hanfei Yu, Bei Ouyang, Shwai He +2LLM Inference EfficiencyExpert Load Balancing

  31. Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving

    Feb 27, 2026Ferran Agullo, Joan Oliveras, Chen Wang +5LLM Inference EfficiencyLLM Serving

  32. More Bang for the Buck: Improving the Inference of Large Language Models at a Fixed Budget using Reset and Discard (ReD)

    Jan 29, 2026Sagi Meir, Tommer D. Keidar, Noam Levi +2LLM Inference EfficiencyLLM Evaluation

  33. Nalar: Workflow-Aware Management of Agentic Applications

    Jan 8, 2026Saurabh Agarwal, Marco Laju, Donghyun Son +4LLM Inference EfficiencyLLM Agent Orchestration

  34. Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring

    Dec 16, 2025Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo +1LLM Inference EfficiencyEfficient Language Model Reasoning

  35. Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

    Nov 11, 2025Jon Saad-Falcon, Avanika Narayan, Hakki Orhun Akengin +12LLM Inference EfficiencyAI Accelerator Inference

  36. Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm

    Sep 28, 2025Kaisen Yang, Tinghe Zhang, Rushi Shah +4LLM Inference EfficiencyLLM Planning

  37. EquivPruner: Boosting Efficiency and Quality in LLM-Based Search via Action Pruning

    May 22, 2025Jiawei Liu, Qisi Chen, Jianshu Zhang +2LLM Inference EfficiencySemantic Textual Similarity

  38. Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models

    Mar 31, 2025Rui Wang, Hongru Wang, Boyang Xue +9LLM Inference EfficiencyCost-Aware Inference

  39. Saving GPU Hours in LLM Inference System Development and Online Workloads with Simulation and DBMS-Inspired Cache Replacement Policies

    Nov 12, 2024Kyoungmin Kim, Jiacheng Li, Kijae Hong +3LLM Inference EfficiencyLLM Inference

  40. UltraQuant: 4-bit KV Caching for Context-Heavy Agents

    Date pendingInesh Chakrabarti, David Limpus, Aditi Ghai Rana +4LLM Inference EfficiencyGPU Kernel Optimization