cs.AIOct 2, 2026

iS-KV: Online Low-Rank KV Cache Compression via Block-Incremental SVD

Authors: Yiren Zhao, Guanghui Song, Tianrui Qin, Kejiang Ye, Cheng-zhong Xu, Xitong Gao

Abstract

Long chain-of-thought reasoning substantially increases KV-cache memory during autoregressive decoding, as every generated token introduces new key and value states and causes the cache to grow linearly with decoding length. Existing KV-cache compression methods typically control this growth through token eviction, but irreversible deletion can remove historical states that later reasoning may need to revisit. SVD-based low-rank compression provides an alternative by retaining all positions with a more compact representation. However, extending it from a fixed prompt cache to online decoding is non-trivial. Through our investigation, we find that if the basis is updated for new tokens while old tokens keep their coordinates in the old basis, the stored history drifts substantially. Based on this observation, we propose iS-KV, an online low-rank KV-cache compression method for long-horizon reasoning. iS-KV keeps a recent window exact while incrementally folding older states into bounded-rank representations. As the low-rank basis evolves, it synchronizes historical coordinates with the updated basis to maintain representation consistency. On DeepSeek-R1-Distill-Llama-8B, iS-KV achieves 82.6% accuracy at 4.06-fold persistent-KV compression, close to the original model's 83.6%. On Qwen3-8B, it achieves 89.2% accuracy at 5.64-fold compression. Under matched memory budgets, iS-KV consistently outperforms token-eviction baselines.

Explore similar work

CardsList
  1. PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression

    Aug 24, 2026Zizhong Wang, Jieying Wang, Zhao Zhang +1

  2. S4^4R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching

    Aug 1, 2026Jialong Han, You Wu, Kewei TuKV CachingLong-Context Language Model Inference

  3. AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD

    Oct 3, 2026Sara Abdali, Jongwoo Ko, Pashmina Cameron