LLM Inference Efficiency

LLM: Large Language Model

Latest papers 234

All topics
CardsList
  1. Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning

    Apr 19, 2026Raman Saparkhan, Majd Hawasly, Md Rizwan Parvez +1LLM Inference EfficiencySelf-Consistency Decoding

  2. Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling

    Apr 19, 2026Zizhang Luo, Yuhao Luo, Youwei Xiao +3LLM Inference EfficiencyMulti-Agent LLM Systems

  3. HiveMind: OS-Inspired Scheduling for Concurrent LLM Agent Workloads

    Apr 18, 2026Justice Owusu Agyemang, Jerry John Kponyo, Obed Kwasi Somuah +3LLM Inference EfficiencyAI Coding Agents

  4. When Agents Go Quiet: Output Generation Capacity and Format-Cost Separation for LLM Document Synthesis

    Apr 17, 2026Justice Owusu Agyemang, Michael Agyare, Miriam Kobbinah +2LLM Inference EfficiencyLong-Form Document Generation

  5. PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving

    Apr 14, 2026Xu Bai, Muhammed Tawfiqul Islam, Chen Wang +1LLM Inference EfficiencyLLM Serving

  6. When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models

    Mar 27, 2026Juan Gabriel Kostelec, Qinghai GuoLLM Inference EfficiencyLLM Evaluation

  7. Clinical Note Bloat Reduction for Efficient LLM Use

    Mar 21, 2026Jordan L. Cahoon, Chloe Stanwyck, Asad Aali +6LLM Inference EfficiencyElectronic Health Records

  8. MoEless: Efficient MoE LLM Serving with Serverless Experts

    Mar 6, 2026Hanfei Yu, Bei Ouyang, Shwai He +2LLM Inference EfficiencyExpert Load Balancing

  9. Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving

    Feb 27, 2026Ferran Agullo, Joan Oliveras, Chen Wang +5LLM Inference EfficiencyLLM Serving

  10. More Bang for the Buck: Improving the Inference of Large Language Models at a Fixed Budget using Reset and Discard (ReD)

    Jan 29, 2026Sagi Meir, Tommer D. Keidar, Noam Levi +2LLM Inference EfficiencyLLM Evaluation

  11. Nalar: Workflow-Aware Management of Agentic Applications

    Jan 8, 2026Saurabh Agarwal, Marco Laju, Donghyun Son +4LLM Inference EfficiencyLLM Agent Orchestration

  12. Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring

    Dec 16, 2025Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo +1LLM Inference EfficiencyEfficient Language Model Reasoning

  13. Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

    Nov 11, 2025Jon Saad-Falcon, Avanika Narayan, Hakki Orhun Akengin +12LLM Inference EfficiencyAI Accelerator Inference

  14. Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm

    Sep 28, 2025Kaisen Yang, Tinghe Zhang, Rushi Shah +4LLM Inference EfficiencyLLM Planning

  15. EquivPruner: Boosting Efficiency and Quality in LLM-Based Search via Action Pruning

    May 22, 2025Jiawei Liu, Qisi Chen, Jianshu Zhang +2LLM Inference EfficiencySemantic Textual Similarity

  16. Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models

    Mar 31, 2025Rui Wang, Hongru Wang, Boyang Xue +9LLM Inference EfficiencyCost-Aware Inference

  17. Saving GPU Hours in LLM Inference System Development and Online Workloads with Simulation and DBMS-Inspired Cache Replacement Policies

    Nov 12, 2024Kyoungmin Kim, Jiacheng Li, Kijae Hong +3LLM Inference EfficiencyLLM Inference

  18. UltraQuant: 4-bit KV Caching for Context-Heavy Agents

    Date pendingInesh Chakrabarti, David Limpus, Aditi Ghai Rana +4LLM Inference EfficiencyGPU Kernel Optimization