KV-Cache Offloading

Momentum

5 papers in the last four weeks, with none the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 23

All topics
CardsList
  1. TempoKV: Timely Staging of LLM KV Caches for Memory-Semantic Flash

    Sep 28, 2026Jay H. Park, Hyungjun Kim, Dong KimKV-Cache OffloadingKV-Cache Management

  2. PulseInfer: I/O-Centric Sparse KV Cache Offloading for Efficient Long-Context LLM Decoding

    Sep 28, 2026Qiuyang Zhang, Kai Zhou, Kai Lu +6KV-Cache OffloadingLong-Context Language Model Inference

  3. EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?

    Sep 27, 2026Kunming Shao, Jierun Chen, Jiangnan Yu +7KV-Cache OffloadingKV-Cache Management

  4. Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs

    Sep 11, 2026Joseph Kanichai, Tiziano De Matteis, Animesh TrivediKV-Cache OffloadingKV Caching

  5. Long-Context Fine-Tuning with Limited VRAM

    Jul 16, 2026Vladimir Fedosov, Aleksandr Sazhin, Artemiy Grinenko +1Fine-TuningKV-Cache Offloading

  6. SeKV: Resolution-Adaptive KV Cache with Hierarchical Semantic Memory for Long-Context LLM Inference

    Jun 30, 2026Amirhossein Abaskohi, Giuseppe Carenini, Peter West +1KV-Cache OffloadingKV Caching

  7. Hierarchical Global Attention (HGA)

    Jun 29, 2026Woernle Frank, Fedosov Vladimir, Grinenko ArtemiyKV-Cache OffloadingLong-Context Modeling

  8. HERALD: High-Throughput Block Diffusion LLM Serving via CPU-GPU Cooperative KV Cache Retrieval

    Jun 19, 2026Omin Kwon, Doyeon Kim, Jongseok Park +3KV-Cache OffloadingDiffusion Model Serving

  9. SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference

    Jun 3, 2026Yaosheng Fu, Guangxuan Xiao, Xin Dong +2KV-Cache OffloadingLong-Context Language Model Inference

  10. KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference

    May 18, 2026Jian Lin, Jiazhi Mi, Zicong Hong +5KV-Cache OffloadingKV Caching

  11. VeriCache: Turning Lossy KV Cache into Lossless LLM Inference

    May 17, 2026Jiayi Yao, Samuel Shen, Kuntai Du +7KV-Cache OffloadingKV Caching

  12. ObjectCache: Layerwise Object-Storage Retrieval for KV Cache Reuse

    May 16, 2026Yu Zhu, Aditya Dhakal, Yunming Xiao +2LLM Inference EfficiencyKV-Cache Offloading

  13. Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning

    May 10, 2026Aojie Yuan, Tianqi Shen, Dajun ZhangKV-Cache OffloadingKV Caching

  14. Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving

    Apr 29, 2026Zihan Zhao, Baotong Lu, Shengjie Lin +8LLM ServingKV-Cache Offloading

  15. DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference

    Apr 29, 2026Bodon Jeong, Hongsu Byun, Youngjae Kim +4On-Device Language Model InferenceKV-Cache Offloading

  16. MTServe: Efficient Serving for Generative Recommendation Models with Hierarchical Caches

    Apr 24, 2026Xin Wang, Chi Ma, Shaobin Chen +14KV-Cache OffloadingKV Caching

  17. SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference

    Apr 23, 2026Hongyao Liu, Liuqun Zhai, Junyi Wang +3On-Device Language Model InferenceKV-Cache Offloading

  18. KV Cache Offloading for Context-Intensive Tasks

    Apr 9, 2026Andrey Bocharnikov, Ivan Ermakov, Denis Kuznedelev +2KV-Cache OffloadingKV Caching

  19. NOSA: Native and Offloadable Sparse Attention

    Oct 15, 2025Yuxiang Huang, Pengjie Wang, Jicheng Han +9KV-Cache OffloadingGPU Acceleration

  20. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

    Jun 24, 2024Ruoyu Qin, Zheming Li, Weiran He +4LLM ServingDisaggregated LLM Serving