LLM Serving

LLM: Large Language Model

Momentum

17 papers in the last four weeks, up 143% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 144

All topics
CardsList
  1. Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving

    Apr 29, 2026Zihan Zhao, Baotong Lu, Shengjie Lin +8LLM ServingKV-Cache Offloading

  2. Pythia: Exploiting Workflow Predictability for Efficient Agent-Native LLM Serving

    Apr 28, 2026Shan Yu, Junyi Shu, Yuanjiang Ni +14LLM Inference EfficiencyMulti-Agent LLM Systems

  3. Latency and Cost of Multi-Agent Intelligent Tutoring at Scale

    Apr 27, 2026Iizalaarab Elhaimeur, Nikos ChrisochoidesLLM Inference EfficiencyMulti-Agent LLM Systems

  4. RouteNLP: Closed-Loop LLM Routing with Conformal Cascading and Distillation Co-Optimization

    Apr 26, 2026Dongxin Guo, Jikun Wu, Siu Ming YiuLLM ServingCost-Aware Inference

  5. Continuous Semantic Caching for Low-Cost LLM Serving

    Apr 21, 2026Baran Atalar, Xutong Liu, Jinhang Zuo +3LLM ServingCost-Aware Inference

  6. SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving

    Apr 19, 2026Christian LysenstøenLLM Inference EfficiencyLLM Serving

  7. Bit-Flip Vulnerability of Shared KV-Cache Blocks in LLM Serving Systems

    Apr 19, 2026Yuji Yamamoto, Satoshi MatsuuraLLM ServingAdversarial Attacks on LLMs

  8. PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving

    Apr 14, 2026Xu Bai, Muhammed Tawfiqul Islam, Chen Wang +1LLM Inference EfficiencyLLM Serving

  9. FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation

    Mar 10, 2026Yinpeng Wu, Yitong Chen, Lixiang Wang +3LLM ServingOn-Device Language Model Inference

  10. MoEless: Efficient MoE LLM Serving with Serverless Experts

    Mar 6, 2026Hanfei Yu, Bei Ouyang, Shwai He +2LLM Inference EfficiencyExpert Load Balancing

  11. Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving

    Feb 27, 2026Ferran Agullo, Joan Oliveras, Chen Wang +5LLM Inference EfficiencyLLM Serving

  12. Learning to Route and Schedule LLMs from User Retrials via Contextual Queueing Bandits

    Feb 2, 2026Seoungbin Bae, Junyoung Son, Dabeen LeeLLM ServingLLM Routing

  13. xGR: Efficient Generative Recommendation Serving at Scale

    Dec 12, 2025Qingxiao Sun, Tongxuan Liu, Shen Zhang +13LLM ServingLLM Inference Acceleration

  14. Systematic Exploration of Multi-core Architectures for Efficient LLM Serving using WaferAI-SIM

    Oct 7, 2025Tianhao Zhu, Dahu Feng, Erhu Feng +1AI Accelerator InferenceLLM Serving

  15. Ranking Before Serving: Low-Latency LLM Serving via Pairwise Learning-to-Rank

    Sep 25, 2025Yiheng Tao, Yihe Zhang, Matthew Dearing +4LLM ServingLearning to Rank

  16. VoltanaLLM: Energy-Efficient and SLO-Aware Disaggregated LLM Serving via Adaptive Frequency Control and State-Space Routing

    Sep 5, 2025Jiahuan Yu, Aryan Taneja, Junfeng Lin +1LLM ServingEnergy-Efficient ML

  17. LLM Serving Optimization with Variable Prefill and Decode Lengths

    Aug 8, 2025Meixuan Wang, Yinyu Ye, Zijie ZhouLLM ServingLLM Inference Scheduling

  18. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

    Jun 24, 2024Ruoyu Qin, Zheming Li, Weiran He +4LLM ServingDisaggregated LLM Serving

  19. OUTLETS: Output-Length Prediction from Speculative Decoding Backbones

    Date pendingWeihuang Wen, Yingying Liu, Yichuan Liu +5LLM ServingLLM Inference Scheduling