cs.LGSep 28, 2026

Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression

Authors: Xingyu Zhu, Pu, Yi, Ziheng Cheng, Ang Lv, Jing Liu, Lexing Ying, Yiyuan Ma, +1 more

Organizations: Luke · ByteDance Seed · University of California, Berkeley

Abstract

Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position relative to compression-window boundaries. We uncover a systematic asymmetry in models using such compression: the same information can be easy to retrieve at one phase and difficult at another. We call this periodic variation in retrieval performance phase sensitivity. In large open-weight models with such compression, long-context retrieval accuracy can differ by up to 40 percentage points across phases, revealing periodic weak spots that average benchmark scores can conceal. To investigate this behavior, we pretrain a family of transformers from scratch across multiple KV-compression designs, reproducing phase sensitivity across the variants. Mechanistic analysis using causal interventions in these models reveals phase specialization: different attention components contribute asymmetrically to retrieving information at different source phases. We further analyze idealized retrieval models, showing how gradient flow dynamics may favor sharp phase specialization. Evaluating models with chunked KV-cache compression thus requires measuring across compression phases: high average accuracy can coexist with systematic positional failures.

Explore similar work

CardsList
  1. FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference

    Jul 7, 2026Anna Córdoba, Adam Puente Tercero, Nerea Angulo Hijo +4Key-Value Cache CompressionDepthweave-Kv

  2. DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression

    Jul 7, 2026Anna Cordoba, Adam Puente Tercero, Nerea Angulo Hijo +4Key-Value Cache CompressionDepthweave-Kv

  3. Training Transformers for KV Cache Compressibility

    May 7, 2026Yoav Gelberg, Yam Eitan, Michael Bronstein +2Key-Value Cache CompressionEfficient Long-Context Inference