LLM Serving

LLM: Large Language Model

Momentum

17 papers in the last four weeks, up 143% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 142

All topics
CardsList
  1. Pythia: Exploiting Workflow Predictability for Efficient Agent-Native LLM Serving

    Apr 28, 2026Shan Yu, Junyi Shu, Yuanjiang Ni +14LLM Inference EfficiencyMulti-Agent LLM Systems

  2. Latency and Cost of Multi-Agent Intelligent Tutoring at Scale

    Apr 27, 2026Iizalaarab Elhaimeur, Nikos ChrisochoidesLLM Inference EfficiencyMulti-Agent LLM Systems

  3. RouteNLP: Closed-Loop LLM Routing with Conformal Cascading and Distillation Co-Optimization

    Apr 26, 2026Dongxin Guo, Jikun Wu, Siu Ming YiuLLM ServingCost-Aware Inference

  4. Continuous Semantic Caching for Low-Cost LLM Serving

    Apr 21, 2026Baran Atalar, Xutong Liu, Jinhang Zuo +3LLM ServingCost-Aware Inference

  5. SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving

    Apr 19, 2026Christian LysenstøenLLM Inference EfficiencyLLM Serving

  6. Bit-Flip Vulnerability of Shared KV-Cache Blocks in LLM Serving Systems

    Apr 19, 2026Yuji Yamamoto, Satoshi MatsuuraLLM ServingAdversarial Attacks on LLMs

  7. PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving

    Apr 14, 2026Xu Bai, Muhammed Tawfiqul Islam, Chen Wang +1LLM Inference EfficiencyLLM Serving

  8. FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation

    Mar 10, 2026Yinpeng Wu, Yitong Chen, Lixiang Wang +3LLM ServingOn-Device Language Model Inference

  9. MoEless: Efficient MoE LLM Serving with Serverless Experts

    Mar 6, 2026Hanfei Yu, Bei Ouyang, Shwai He +2LLM Inference EfficiencyExpert Load Balancing

  10. Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving

    Feb 27, 2026Ferran Agullo, Joan Oliveras, Chen Wang +5LLM Inference EfficiencyLLM Serving

  11. Learning to Route and Schedule LLMs from User Retrials via Contextual Queueing Bandits

    Feb 2, 2026Seoungbin Bae, Junyoung Son, Dabeen LeeLLM ServingLLM Routing

  12. xGR: Efficient Generative Recommendation Serving at Scale

    Dec 12, 2025Qingxiao Sun, Tongxuan Liu, Shen Zhang +13LLM ServingLLM Inference Acceleration

  13. Systematic Exploration of Multi-core Architectures for Efficient LLM Serving using WaferAI-SIM

    Oct 7, 2025Tianhao Zhu, Dahu Feng, Erhu Feng +1AI Accelerator InferenceLLM Serving

  14. Ranking Before Serving: Low-Latency LLM Serving via Pairwise Learning-to-Rank

    Sep 25, 2025Yiheng Tao, Yihe Zhang, Matthew Dearing +4LLM ServingLearning to Rank

  15. VoltanaLLM: Energy-Efficient and SLO-Aware Disaggregated LLM Serving via Adaptive Frequency Control and State-Space Routing

    Sep 5, 2025Jiahuan Yu, Aryan Taneja, Junfeng Lin +1LLM ServingEnergy-Efficient ML

  16. LLM Serving Optimization with Variable Prefill and Decode Lengths

    Aug 8, 2025Meixuan Wang, Yinyu Ye, Zijie ZhouLLM ServingLLM Inference Scheduling

  17. OUTLETS: Output-Length Prediction from Speculative Decoding Backbones

    Date pendingWeihuang Wen, Yingying Liu, Yichuan Liu +5LLM ServingLLM Inference Scheduling