Large Language Model Serving

Recent momentum

+9%

12 papers in the last 28 days · 0.2% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

4 new papers

A weekly snapshot of new work published in Large Language Model Serving.

Period ending 2026-09-14

5 new papers

A weekly snapshot of new work published in Large Language Model Serving.

Period ending 2026-09-07

4 new papers

A weekly snapshot of new work published in Large Language Model Serving.

104 papers

Latest in Large Language Model Serving

Open your feed →
CardsList
  1. AgentKV: Phase-Aware KV Eviction for Agentic LLMs

    Sep 14, 2026Taowen Tony Liu, Jeffrey T. H. Wong, Can Xiao +3Key-Value Cache EvictionKey-Value

  2. LaCache: Robust Semantic Caching for LLM Serving

    Aug 3, 2026Jiacheng Liang, Yuhui Wang, Tanqiu Jiang +1CacheLanguage Model Evasion Attacks

  3. Kalypso: Relational LLM Serving

    Jul 26, 2026Hojae Son, Md Ashraful Islam, Huy Gia Cao +2Large Language Model ServingLarge Language

  4. Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving

    Jul 2, 2026Shrikara Arun, Anjaly Parayil, Srikant Bharadwaj +2PrefillPer-Token Latency

  5. Cluster, Route, Escalate: Cascaded Framework for Cost-Aware LLM Serving

    Jun 25, 2026Yasmin Moslem, Magdalena Kacmajor, Vasudevan Nedumpozhimana +11Large Language Model ServingEscalation