LLM Inference Serving

LLM: Large Language Model

Momentum

6 papers in the last four weeks, against 1 the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 51

All topics
CardsList
  1. Frontier: Towards Comprehensive and Accurate LLM Inference Simulation

    May 20, 2026Yicheng Feng, Xin Tan, Yangtao Deng +3LLM ServingDisaggregated LLM Serving

  2. KVBuffer: IO-aware Serving for Linear Attention

    May 18, 2026Longwei Zou, Lin ZhongLLM Inference ServingLinear Attention

  3. Enabling Performant and Flexible Model-Internal Observability for LLM Inference

    May 11, 2026Nengneng Yu, Sixian Xiong, Yibo Zhao +2LLM Inference ServingEfficient Language Model Inference

  4. KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving

    May 10, 2026Zhiqing Zhong, Zhijing Ye, Jian Zhang +3KV CachingKV-Cache Management

  5. Regulating Branch Parallelism in LLM Serving

    May 7, 2026Swapnil Gandhi, Siva Hari, William J. Dally +1LLM ServingLLM Inference Scheduling

  6. Towards Distributed Inference of LLMs on a P2P Network

    May 7, 2026Shabari S Nair, Krishanu SainiLLM InferenceKV Caching

  7. Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving

    May 7, 2026Bole Ma, Jan Eitzinger, Harald KöstlerKV CachingLLM Inference Serving

  8. A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints

    May 6, 2026Chengyi Nie, Nian Si, Zijie ZhouLLM ServingLLM Inference

  9. Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference

    May 1, 2026Yuxuan Gao, Megan Wang, Yi Ling YuLLM EvaluationEnergy-Efficient ML

  10. FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving

    Apr 29, 2026Minghe Wang, Trever Schirmer, Mohammadreza Malekabbasi +1Mixture-of-Experts Language ModelsLLM Serving

  11. Scalable Inference Architectures for Compound AI Systems: A Production Deployment Study

    Apr 28, 2026Srikanta Prasad S, Utkarsh AroraLLM Inference Serving

  12. KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving

    Apr 17, 2026Yichao Yuan, Mosharaf Chowdhury, Nishil TalatiEfficient InferenceLLM Inference Serving

  13. Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines

    Apr 16, 2026Otto White, Marcel Wagenländer, Britannio Jarrett +7LLM InferenceLLM Inference Scheduling

  14. ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System

    Feb 14, 2026Hao Kang, Ziyang Li, Weili Xu +7LLM Agent OrchestrationLLM Inference Scheduling

  15. Ranking Before Serving: Low-Latency LLM Serving via Pairwise Learning-to-Rank

    Sep 25, 2025Yiheng Tao, Yihe Zhang, Matthew Dearing +4LLM ServingLearning to Rank

  16. Saving GPU Hours in LLM Inference System Development and Online Workloads with Simulation and DBMS-Inspired Cache Replacement Policies

    Nov 12, 2024Kyoungmin Kim, Jiacheng Li, Kijae Hong +3LLM Inference EfficiencyLLM Inference