cs.LGOct 8, 2026

Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration

Authors: Kaicheng Xiao, Liran Dong, Haotian Li, Guoliang Xing

Organizations: The Chinese University of Hong Kong

Abstract

KV-cache quantization and linear attention are two representative approaches to tackling the storage and computational costs of Transformers. KV-cache quantization compresses individual KV entries into discrete codes but retains all entries, whereas linear attention recurrently aggregates multiple historical KV contributions into a fixed-size continuous state but can introduce interference. This contrast raises the question of whether per-KV compression and multi-KV aggregation can be bridged within a single mechanism for efficient attention. We identify RAM-Net as such a bridge through soft assignments over a discrete address space. These assignments determine recurrent updates to the continuous slot state associated with each address. Under a restricted RAM-Net construction, we prove that soft address assignments extend hard quantized matching to a separable read-write overlap that locally approximates full-attention similarity and supports recurrent aggregation. These connections further enable Transformer-to-RAM-Net weight migration through a new path based on a soft-quantized intermediate construction. Across nine pretrained Transformer models from 0.3B to 7B parameters, RAM-Net recovers an average of 87.1% of the teachers' accuracy gains over random guessing across six commonsense and knowledge tasks using only a 500M-token budget per model.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ResidualQuant: KV Cache Quantization for Looped Transformers with 2-Bit Residuals

    Oct 7, 2026Heejun Kim, Junyoung Lee, SangLyul Cho +3KV-Cache QuantizationEfficient Transformer Inference

  2. A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal Attention

    Oct 6, 2026Davis Wertheimer, Haochen Shen, Ahan Gupta +6KV-Cache CompressionTransformer Attention

  3. ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression

    Jul 31, 2026Yuhang Zhan, Lisi Chen, Shuo ShangSoftmax AttentionKV Caching