KV-cache quantization and linear attention are two representative approaches to tackling the storage and computational costs of Transformers. KV-cache quantization compresses individual KV entries into discrete codes but retains all entries, whereas linear attention recurrently aggregates multiple historical KV contributions into a fixed-size continuous state but can introduce interference. This contrast raises the question of whether per-KV compression and multi-KV aggregation can be bridged within a single mechanism for efficient attention. We identify RAM-Net as such a bridge through soft assignments over a discrete address space. These assignments determine recurrent updates to the continuous slot state associated with each address. Under a restricted RAM-Net construction, we prove that soft address assignments extend hard quantized matching to a separable read-write overlap that locally approximates full-attention similarity and supports recurrent aggregation. These connections further enable Transformer-to-RAM-Net weight migration through a new path based on a soft-quantized intermediate construction. Across nine pretrained Transformer models from 0.3B to 7B parameters, RAM-Net recovers an average of 87.1% of the teachers' accuracy gains over random guessing across six commonsense and knowledge tasks using only a 500M-token budget per model.
Figures & tables
Figure 1: Comparison of attention predictors in QK space. For each sub-plot, the left shows how historical KV information is stored or aggregated, while the right visualizes the resulting ft(q) over QK space. RAM-Net retains the discrete-code representation of KV-cache quantization and the continuous-state aggregation of linear attention.
Figure 2: Overview of Sec. 3 . RAM-Net connects discrete-code Soft-NN and linear predictors through the separable overlap between read and write distributions built on circular soft addressing, which locally approximates Soft-NN matching.
Family
Teacher
Model
Tokens
Ctx.
Wiki PPL ↓
ARC-E ↑
ARC-C ↑
PIQA ↑
SciQ ↑
COPA ↑
Hella. ↑
Avg.↑
Avg.R↑
FLA
transformer-340M- 10B
Teacher
10B
4K
30.05
57.4
24.1
65.6
82.5
73.0
32.1
55.8
86.5
RAM-Net
0.5B
2K
35.14
54.7
23.7
64.7
80.2
65.0
31.0
53.2
transformer-1.3B- 100B
Teacher
100B
2K
17.66
56.1
23.9
70.1
85.3
75.0
38.5
58.1
86.4
RAM-Net
0.5B
2K
21.93
52.7
24.8
68.3
78.8
71.0
35.6
55.2
transformer-2.7B- 100B
Teacher
100B
2K
15.23
59.4
26.5
71.9
87.9
75.0
42.0
60.4
89.2
RAM-Net
0.5B
2K
19.82
56.5
24.6
70.3
79.9
74.0
38.3
57.3
Table 1: Zero-shot language modeling and commonsense reasoning performance of Teacher-to-RAM-Net migration.
Figure 3: Attn. JS under different circular assignment concentrations β .
Method
FLA-0.3B
FLA-1.3B
FLA-2.7B
Qwen-0.5B
Qwen-1.5B
Qwen-3B
Red.-3B
Mist.-7B
LLaMA-7B
r/w projection
0.561
0.563
0.568
0.625
0.617
0.601
0.611
0.619
0.629
QK projection adapter
0.570
0.558
0.571
0.627
0.629
0.623
0.601
0.615
0.615
Random codebook
0.441
0.440
0.437
0.486
0.476
0.498
0.443
0.430
0.389
Orth. Rand. codebook
0.487
0.515
0.525
0.574
0.566
0.560
0.563
0.568
0.586
Shared k -means codebook
0.415
0.412
0.419
0.466
0.453
0.481
0.421
0.411
0.377
Per-factor PQ codebook
0.412
0.423
0.431
0.487
0.476
0.487
0.446
0.445
0.407
Table 2: Attention JS divergence of intermediate attention construction.
Table 6
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Value fields under different RAM-Net configurations. Keys and queries lie on S1×S1⊂R4 , with the plotted axes representing the angular coordinates of the two factors. Marker colors encode the stored values, while the background shows the predicted value ft(q) . Full attention provides the reference on the left. Rows vary the codeword norm around the reference setting β⋆ , and columns vary Top- K , dp , and mass-aware readout.
Family
Teacher
Layer
Hidden size
Head
KV Head
dhead
Attn.
Native ctx.
FLA
transformer-340M-10B
24
1024
32
32
32
MHA
4K
transformer-1.3B-100B
24
2048
32
32
64
MHA
2K
transformer-2.7B-100B
32
2560
20
20
128
MHA
2K
Qwen2.5
0.5B-Base
24
896
14
2
64
GQA
32K
1.5B-Base
28
1536
12
2
128
GQA
32K
3B-Base
36
2048
16
2
128
GQA
32K
Appendix
Table 5: Architectures of the teacher models.
Phase
Stage
Main Loss
Trainable Parameters
Tokens
Warmup
LR
Intermediate
Low-rank & positional alignment
Attention KL
QK adapters, log-scales
7.209M
0.360M
3×10−4
Circular soft quantization
–
–
0
–
–
Recovery
Head-level
Head-output NMSE
Wr , Ww , scales, γt
49.152M
1.638M
1×10−4
Attn.-block local
Residual-update NMSE
Attention branch
61.440M
2.048M
1×10−4
Attn.-block stack
Residual-update NMSE
Attention branch
36.864M
1.638M
3×10−5
Full-model
0.25 Logit KL + CE
All parameters
345.375M
0.819M
1×10−5
Appendix
Table 6: Overview of Transformer-to-RAM-Net migration. The total training budget is 500M tokens per teacher.
Family
Teacher
Wiki PPL ↓
ARC-E ↑
ARC-C ↑
PIQA ↑
SciQ ↑
COPA ↑
Hella. ↑
Avg. ↑
R ↑
FLA
transformer-340M-10B
35.28 (+0.1)
55.8 (+1.1)
23.9 (+0.2)
64.3 (-0.4)
80.3 (+0.1)
66.0 (+1.0)
31.1 (+0.1)
53.6 (+0.4)
87.7 (+1.2)
transformer-1.3B-100B
21.83 (-0.1)
52.3 (-0.4)
23.7 (-1.1)
67.7 (-0.6)
78.0 (-0.8)
72.0 (+1.0)
35.8 (+0.2)
54.9 (-0.3)
86.4 (-0.1)
transformer-2.7B-100B
20.09 (+0.3)
56.8 (+0.3)
25.4 (+0.8)
69.6 (-0.7)
79.4 (-0.5)
76.0 (+2.0)
37.8 (-0.5)
57.5 (+0.2)
89.6 (+0.4)
Qwen2.5
0.5B-Base
23.55 (+0.2)
64.1 (+0.5)
28.8 (+2.4)
68.4 (-0.1)
81.9 (0.0)
69.0 (-4.0)
36.8 (-0.1)
58.2 (-0.2)
86.1 (+6.9)
1.5B-Base
17.30 (+1.0)
74.8 (-0.5)
40.2 (-1.3)
74.7 (+0.4)
88.4 (-1.4)
77.0 (-5.0)
44.7 (-0.9)
66.6 (-1.5)
90.2 (-4.8)
3B-Base
13.89 (+0.3)
77.2 (0.0)
44.2 (+0.5)
77.0 (0.0)
91.6 (-1.3)
81.0 (+2.0)
50.3 (-0.4)
70.2 (+0.1)
92.3 (+0.8)
Appendix
Table 7: Mass-aware readout results. Values in parentheses indicate the score differences relative to RAM-Net in Table 1 .
Figure 5: Head NMSE under different circular assignment concentrations β .
Readout rule
FLA-0.3B
FLA-1.3B
FLA-2.7B
Qwen-0.5B
Qwen-1.5B
Qwen-3B
Red.-3B
Mist.-7B
LLaMA-7B
w/o
0.277
0.286
0.308
0.292
0.293
0.312
0.312
0.283
0.253
w/
0.259
0.273
0.300
0.290
0.297
0.311
0.319
0.280
0.257
Appendix
Table 8: JS divergence with and without mass-aware readout after enabling GSU.
Wiki PPL ↓
ARC-E ↑
ARC-C ↑
PIQA ↑
SciQ ↑
COPA ↑
Hella. ↑
Avg. ↑
w/o
26.89
59.1
24.7
66.9
82.5
67.0
33.4
55.6
w/
27.10
59.1
26.1
67.5
82.1
73.0
33.3
56.8
Appendix
Table 9: Train-from-scratch results for 340M models with and without mass-aware readout.
Metric
Method
FLA-0.3B
FLA-1.3B
FLA-2.7B
Qwen-0.5B
Qwen-1.5B
Qwen-3B
Red.-3B
Mist.-7B
LLaMA-7B
LCC ↑
Teacher
17.0
48.3
50.7
10.0
14.7
14.1
44.7
70.0
66.7
RAM-Net
19.6
36.4
36.4
18.5
12.7
15.5
32.7
33.4
24.8
MultiNews ↑
Teacher
10.5
6.6
6.3
21.4
25.2
24.5
20.1
15.5
13.8
RAM-Net
9.6
7.2
8.4
12.9
13.9
15.9
8.6
12.1
11.5
SAMSum ↑
Teacher
6.6
6.2
4.9
37.3
42.4
45.5
0.0
28.2
31.4
RAM-Net
8.0
9.3
10.3
20.3
28.7
29.5
0.1
22.1
23.9
Appendix
Table 10: LongBench results for all teacher and RAM-Net pairs.
Figure 6: Effect of dp and order U on intermediate forward KL and head NMSE.
Order
Low-rank KL ↓
Quant. JS ↓
Head NMSE ↓
Logit KL ↓
Wiki PPL ↓
ARC-E ↑
ARC-C ↑
PIQA ↑
SciQ ↑
COPA ↑
Hella. ↑
Avg. ↑
R ↑
4 (256 slots)
1.957
0.253
0.177
0.500
16.58
73.8
38.9
73.6
87.4
82.0
45.6
66.9
90.8
5 (1024 slots)
1.748
0.293
0.197
0.465
16.33
75.3
41.5
74.3
89.8
82.0
45.6
68.1
95.0
6 (4096 slots)
1.555
0.339
0.209
0.476
16.01
74.4
41.6
74.0
89.2
77.0
45.4
66.9
91.8
Appendix
Table 11: Effect of order U on Qwen2.5-1.5B. Logit KL is measured after full-model recovery.
Method
Checkpoint
Tokens (M)
Logit KL ↓
Wiki ↓
ARC-E
ARC-C
PIQA
SciQ
COPA
Hella.
Avg. ↑
R ↑
Teacher
Original
–
–
12.60
75.2
41.1
75.6
94.2
83.0
50.2
69.9
–
Complete
Head-Rec. end
56.4
–
77.56
47.1
25.1
64.2
69.9
69.0
35.2
51.7
43.8
Block-Rec. end
154.7
–
19.44
72.1
37.6
73.6
87.2
79.0
44.1
65.6
86.3
Final
500.0
0.164
16.33
75.3
41.5
74.3
89.8
82.0
45.6
68.1
95.0
Full-Model only
Matched head
56.4
0.419
23.90
65.1
31.5
70.2
83.0
77.0
40.7
61.2
71.1
Matched block
154.7
0.278
19.35
70.1
35.7
73.6
87.3
78.0
43.2
64.7
82.6
Appendix
Table 12: Ablation of the staged migration pipeline on Qwen2.5-1.5B.
Figure 7: Head-level recovery NMSE under different low-rank adapter and codebook constructions.
Figure 8: Stage-wise evolution throughout the migration pipeline of FLA-2.7B.
Looped Transformers improve parameter efficiency by repeatedly applying shared Transformer blocks over multiple recurrent loops, increasing computational depth without increasing the parameter count. However, KV cache memory still scales with the number of loops, becoming a key memory bottleneck that limits batch size and inference throughput. KV cache quantization can alleviate this bottleneck, but existing methods often suffer substantial accuracy degradation at aggressive low-precision regimes. We observe that looped Transformers offer a unique opportunity: KV states across loops are highly similar. Based on this observation, we propose ResidualQuant, which uses the final-loop KV states as a reference and represents the remaining loops with low-precision residuals. Our method further combines least-square scaling and rotations applied to the residuals, as well as loop-wise mixed precision, to enable accurate quantization down to INT2 while retaining efficient reconstruction. Across multiple looped Transformer models and mathematical reasoning and code generation benchmarks, ResidualQuant consistently improves the accuracy-memory tradeoff over state-of-the-art rotation-based KV quantization. In particular, our method retains accuracy close to BF16 under mixed-precision settings while reducing theoretical KV storage by 80.7%, achieving up to 13.0% higher accuracy than the rotation-based baseline at the same memory budget. On an RTX 5090, the reduced KV memory traffic improves fixed-batch decode throughput by up to 2.73x, while the smaller memory footprint enables up to 2x larger batches, improving peak throughput by up to 4.15x.
Heejun Kim, Junyoung Lee, SangLyul Cho +3
KAIST · Yonsei University · Seoul National University
The large KV-cache size of modern LLMs creates a barrier to efficient deployment. Recent work has explored replacing attention layers' RoPE positional embeddings with alternative decay-based mechanisms, which can then be used to prune KV-cache during inference. However, these decay functions have limited expressivity, and in practice devolve into sliding-window-like eviction patterns. In this work, we propose a unifying framework for complementary and novel decay mechanisms, capturing complex key statistics and interactions while preserving expressive RoPE embeddings and Softmax attention. The resulting Universal Attention is a highly expressive and end-to-end trainable architecture, whose composite decay mechanism acts as a natural, adaptive pruning criterion, removing tokens that contribute least to attention computation. Experimentally, Universal Attention achieves state-of-the-art 10× compression on natural language and synthetic task data, while improving downstream performance compared to both state-of-the-art baselines and unpruned oracles. It further demonstrates superior long-context generalization with unprecedented 25× compression at length 16k.
Davis Wertheimer, Haochen Shen, Ahan Gupta +6
IBM Research, USA · SSAIL Lab, University of Illinois Urbana-Champaign
KV cache compression is essential for efficient long-context inference. Existing eviction methods permanently discard unselected tokens and consequently remove their aggregate contribution to attention. Merging-based alternatives preserve more information but can perturb retained keys and values that should remain exact. We observe that the information omitted by cache eviction can be formulated as residual statistics in both the numerator and denominator of softmax attention. Based on this observation, we propose ResKV, which divides a fixed KV budget into an exact main cache and a compact residual cache that reconstructs the contribution of omitted tokens. ResKV lets main-cache tokens and residual entries participate in the same softmax normalization, so residual entries restore both attention numerator and denominator mass rather than acting as a post-hoc correction. A construction-time validation proxy determines residual allocation for each layer and KV head, while a decode-time dynamic gate adjusts residual contributions for individual queries. Comprehensive evaluations on LongBench and RULER, covering query-aware and query-agnostic settings, multiple backbones, cache budgets, and representative compression baselines, demonstrate broad improvements under the same retained KV budget while preserving the practical efficiency of compressed decoding, including peak memory usage and long-context decode throughput.
Yuhang Zhan, Lisi Chen, Shuo Shang
University of Electronic Science and Technology of China