Cache

Momentum

31 papers in the last four weeks, up 94% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 196

All topics
CardsList
  1. PatchKV: Weight-Space Compensation of KV Cache

    Sep 30, 2026Chanryeol Lee, Chanhyuk Lee, Yeonwoo Choi +2Key-Value Cache CompressionKey-Value Cache

  2. ARC-KV: Amortizing Anchor Search for Reconstruction-Based KV Cache Compaction

    Sep 29, 2026Zheyu Shen, Guanhua Wang, Dezhan Tu +5Key-Value Cache CompressionCache

  3. RA-CFGCache: From Branch-Level Criteria to Guided-Risk Control under Classifier-Free Guidance

    Sep 29, 2026Yiming Liu, Ben Wan, Tongxuan Liu +5Classifier-Free GuidanceCache

  4. TempoKV: Timely Staging of LLM KV Caches for Memory-Semantic Flash

    Sep 28, 2026Jay H. Park, Hyungjun Kim, Dong KimKv-Cache ManagementCache

  5. KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems

    Sep 28, 2026Hyesung Jeon, Hyeongju Ha, Seoyoung Lee +2Kv-Cache ManagementKey-Value Cache

  6. PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction

    Sep 28, 2026Hyesung Jeon, Hyeongju Ha, Jae-Joon KimKey-Value CacheCache

  7. EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?

    Sep 27, 2026Kunming Shao, Jierun Chen, Jiangnan Yu +7Kv-Cache ManagementOffloading

  8. RelaxKV: Recomputation Guided by the Query with Sparse Context Attention for Efficient KV Cache Reuse

    Sep 27, 2026Ruoling Qi, Yirui Liu, Xuaner Wu +5CacheKey-Value Cache

  9. Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs

    Sep 27, 2026Yirui Liu, Ruoling Qi, Xuaner Wu +6Kimi Delta AttentionCache

  10. In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion

    Sep 26, 2026Yikai Wang, Xiao Han, Mengmeng Xu +9Autoregressive Video Diffusion ModelsCache

  11. Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs

    Sep 26, 2026Vincent-Daniel Yun, Woosang Lim, Haneul Yoo +3PrefillMulti-Agent Large Language Model Systems

  12. When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse

    Sep 24, 2026Yiyu Liu, Minlan Yu, Juncheng YangEvictionCache

  13. KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation

    Sep 21, 2026Sihyeon Ha, Jaeho Lee, Yo-Seb JeonKey-Value Cache CompressionCache

  14. H-Spec: Parallel Speculative Decoding Without a Drafter-Side KV Cache

    Sep 21, 2026Weifan Jiang, Krishna Teja Chitty-Venkata, Megan Flynn +6Speculative DecodingBlock-Diffusion Drafter

  15. PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving

    Sep 17, 2026Omkar Shewale, Deepak Kumar, Divakar Kumar YadavTime-To-First-TokenCache

  16. Fast-Convergent Meta-RL via Gradient-Clustered BS Sampling for Edge Caching

    Sep 14, 2026Farnaz Niknia, Ping WangProximal Policy OptimizationOffline Reinforcement Learning

  17. Fixed State, Long Reach: What a Constant-Size Cache Buys Block Diffusion at Scale

    Sep 14, 2026Vaibhav Singh, Pierre-André Noël, Torsten Scholak +2Diffusion Language ModelsCache

  18. OmniKVQuant: KV Cache Quantization for Omni-LLMs

    Sep 11, 2026Suho Yoo, Hyunjong Ok, Jongmin Choi +2Key-Value Cache QuantizationCache

  19. TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents

    Sep 9, 2026Yuhao Wang, Mu Qiao, Xindong Zhang +3Visual Token PruningTraces

  20. Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both?

    Sep 9, 2026Fumihiko Tachibana, Daisuke Miyashita, Jun DeguchiKey-Value CacheCache

  21. Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning

    Sep 3, 2026Heng Wang, Jielin Qiu, Wenting Zhao +8Key-Value Cache EvictionKey-Value Cache

  22. HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models

    Sep 2, 2026Renjie Xie, Juncheng Yang, Aoting Hu +4Efficient Long-Context InferenceKey-Value Cache

  23. mzCache: On-Device LLM Memory Management under Multitasking

    Sep 1, 2026Hongseung Yu, Minsung Kim, Jongseok Park +1LLM Inference OptimizationLarge Language Model Memory

  24. Invalidation Contracts for Cross-Episode Agent Memory

    Aug 31, 2026Michael Wu, Arquimedes CanedoGradient StalenessContracts

  25. Tail-Replay: Escaping the Curse of Linear Attention in Prefix Caching for Hybrid LLMs

    Aug 31, 2026Yirui Liu, Ruoling Qi, Xuaner Wu +2Kimi Delta AttentionLinear Attention

  26. RegionCache: Semantic-Aware Region Reuse for Efficient Multi-Turn Image Generation

    Aug 30, 2026Peizheng Li, Xin Ai, Hanyuan Liu +2Diffusion-Based Image EditingImage Editing

  27. TwinKV: A Composable Repair Pass for KV Cache Eviction via Pairwise Key Redundancy

    Aug 27, 2026Hong Chen, Yudong Zeng, Yongwei Huang +4Key-Value Cache EvictionCache

  28. StepKV: Step-Aware KV Cache Compression for LLM Agents

    Aug 26, 2026Boyu Feng, Jiahong Liu, Yifan Li +7Key-Value Cache CompressionCache

  29. Beyond Attention Masks: Instruction Anchoring for Efficient In-Context Diffusion Generation

    Aug 21, 2026Yangshuai Liu, Zheming Li, Jiaao Li +4Diffusion TransformersMultimodal Generation

  30. From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion

    Aug 13, 2026Xichen Ye, Yifan Wu, Zhikang Xie +3Diffusion ModelsDenoising Trajectory

  31. Neural Introspection Gating for Adaptive KV-Cache Reuse in Vision-Language-Action Models

    Aug 11, 2026Zhijie Wu, Kento Kawaharazuka, Kei OkadaDiffusion-Based Vision-Language-ActionsCache

  32. BAG: Budget-Aware Gating for Diffusion Caching

    Aug 10, 2026Tong Zhao, Mingkun Lei, Yucheng Han +1Diffusion TransformersCache

  33. RippleKV: Cross-Layer KV Cache Allocation via Perturbation Propagation

    Aug 9, 2026Dongjie Xu, Kai Qian, Julius +6Key-Value Cache CompressionKey-Value Cache

  34. KGCache: Amortized Subgraph Retrieval for KG Reasoning with LLMs

    Aug 8, 2026Uros Stanic, Changcheng Yuan, Sabuj Laskar +1Knowledge GraphsCache

  35. SPECTRA: Pushing the KV Cache Beyond the 2-Bit Cliff via Spectral Transform Coding

    Aug 8, 2026Jiamu Zhang, Liang Wu, Kelly Wan +2Key-Value CacheCache

  36. CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG

    Aug 7, 2026Gyuwan Kim, Cheoneum Park, Tao YangRetrieval-Augmented Generation PipelinesCache

  37. Test-Time Adaptation with Online Personalized Energy-Based Cache for Fine-Grained Video Expression Recognition

    Aug 6, 2026Masoumeh Sharafi, Muhammad Osama Zeeshan, Soufiane Belharbi +3Facial Expression RecognitionTest-Time Adaptation

  38. DG-FedReuse: Proxy-Gradient-Gated Cached-Update Reuse with Matched Sparse Uplink Accounting

    Aug 5, 2026Rahil Aftab, Vineet Kumar Rakesh, Soumya Mazumdar +1Federated LearningFedavg

  39. Caching for the Future: Scrub Jay Episodic Memory Principles for Agent Memory Systems

    Aug 5, 2026Kartikey Singh Bhandari, Aarya Wadhwani, Dhruv Kumar +1Episodic MemoryCache

  40. Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms

    Aug 4, 2026Samuel Fernández-Menduiña, Amir Ziashahabi, Eduardo Pavez +2Key-Value Cache QuantizationResidual Vector Quantization

  41. AnchorKV: Anchor-Residual KV Cache Compression

    Aug 3, 2026Malik Khalaf, Yara Shamshoum, Nitzan Hodos +2Key-Value Cache CompressionKey-Value Cache