Prefix Caching

Momentum

12 papers in the last four weeks, against 1 the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 25

All topics
CardsList
  1. CADOC: Cache-Aware Dynamic Object Context for Long-Horizon Agents

    Sep 29, 2026Junjie Yao, Zhangchen Zhou, Zhi-Qin John XuCost-Aware InferencePrefix Caching

  2. TempoKV: Timely Staging of LLM KV Caches for Memory-Semantic Flash

    Sep 28, 2026Jay H. Park, Hyungjun Kim, Dong KimKV-Cache OffloadingKV-Cache Management

  3. Dynamic Flow, Static Graph: KV Cache Reuse for Efficient LLM Serving on Mobile NPUs

    Sep 28, 2026Zhengxiang Huang, Shengheng Chen, Chaoyue Niu +6AI Accelerator InferenceOn-Device Language Model Inference

  4. EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?

    Sep 27, 2026Kunming Shao, Jierun Chen, Jiangnan Yu +7KV-Cache OffloadingKV-Cache Management

  5. Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs

    Sep 27, 2026Yirui Liu, Ruoling Qi, Xuaner Wu +6Long-Context Language Model InferenceLLM Inference Acceleration

  6. PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving

    Sep 17, 2026Omkar Shewale, Deepak Kumar, Divakar Kumar YadavLLM Inference EfficiencyLLM Inference

  7. Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs

    Sep 11, 2026Joseph Kanichai, Tiziano De Matteis, Animesh TrivediKV-Cache OffloadingKV Caching

  8. KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints

    Sep 9, 2026Xi Shi, Mengxin Zheng, Qian LouKV CachingLLM Inference Acceleration

  9. Tail-Replay: Escaping the Curse of Linear Attention in Prefix Caching for Hybrid LLMs

    Aug 31, 2026Yirui Liu, Ruoling Qi, Xuaner Wu +2LLM Inference AccelerationPrefix Caching

  10. CacheWeaver: Cache-Aware Evidence Ordering for Efficient Grounded RAG Inference

    Jun 18, 2026Kaizhen Tan, Rong Gu, Mingyuan LiLLM Inference EfficiencyRetrieval-Augmented Generation

  11. MiniPIC: Flexible Position-Independent Caching in <100LOC

    Jun 11, 2026Nathan Ordonez, Thomas ParnellLLM Inference EfficiencyKV Caching

  12. Enabling KV Caching of Shared Prefix for Diffusion Language Models

    May 26, 2026Younghun Go, Jaehoon Han, Changyong Shin +2KV CachingKV-Cache Management

  13. ObjectCache: Layerwise Object-Storage Retrieval for KV Cache Reuse

    May 16, 2026Yu Zhu, Aditya Dhakal, Yunming Xiao +2LLM Inference EfficiencyKV-Cache Offloading

  14. Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches

    May 12, 2026Shaoke Fang, Ziang Li, Wenfei Wu +3LLM Inference EfficiencyKV Caching

  15. Agent-X: Full Pipeline Acceleration of On-device AI Agents

    May 11, 2026Jinha Chung, Byeongjun Shin, Jiin Kim +1LLM Inference EfficiencyOn-Device Language Model Inference

  16. PEEK: Predictive Queue-Informed KV Cache Management for LLM Serving

    May 10, 2026Bing Xie, Zhipeng Wang, Masahiro Tanaka +1LLM ServingKV Caching

  17. PRISM: Fast Online LLM Serving via Scheduling-Memory Co-design

    May 9, 2026Xingyu Qu, Tianhao Lin, Yiqi Li +2LLM ServingKV Caching

  18. Towards Distributed Inference of LLMs on a P2P Network

    May 7, 2026Shabari S Nair, Krishanu SainiLLM InferenceKV Caching