cs.CLOct 8, 2026

VFold: Symmetry-Aware Cross-Layer Value Cache Compression

Authors: Neha Verma, Sungwon Kim, Kenton Murray, Kevin Duh

Organizations: Johns Hopkins University · George Mason University

Abstract

While caching key-value (KV) states accelerates Large Language Model (LLM) decoding, this cache can dominate memory usage at long context lengths. One solution is to compress this memory by exploiting inter-layer cache similarities. However, most existing techniques necessitate architectural changes to LLMs and incur substantial overhead. In this work, we propose a symmetry-aware value cache merging strategy that reduces cache memory while avoiding both harmful performance degradation and architectural overhead during decoding. Furthermore, we show that this approach can be exploited alongside existing cache compression techniques, composing with high-ratio quantization or key cache pruning to reach compression ratios that neither method reaches alone, with minimal additional cost. Ultimately, our findings reveal a major source of underutilized capacity in the value cache, offering a simple yet highly effective direction for scaling context windows under memory constraints.

Figures & tables

Explore similar work

CardsList
  1. KV-Kaizen: Learning Context-Adaptive Cache Compression Choices

    Sep 29, 2026Joao Monteiro, Louis Béthune, Anastasiia Filippova +3LLM Inference AccelerationLow-Rank Compression

  2. S4^4R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching

    Aug 1, 2026Jialong Han, You Wu, Kewei TuKV CachingLong-Context Language Model Inference

  3. AnchorKV: Anchor-Residual KV Cache Compression

    Aug 3, 2026Malik Khalaf, Yara Shamshoum, Nitzan Hodos +2KV CachingLong-Context Language Model Inference