Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration
Organizations: The Chinese University of Hong Kong
Abstract
KV-cache quantization and linear attention are two representative approaches to tackling the storage and computational costs of Transformers. KV-cache quantization compresses individual KV entries into discrete codes but retains all entries, whereas linear attention recurrently aggregates multiple historical KV contributions into a fixed-size continuous state but can introduce interference. This contrast raises the question of whether per-KV compression and multi-KV aggregation can be bridged within a single mechanism for efficient attention. We identify RAM-Net as such a bridge through soft assignments over a discrete address space. These assignments determine recurrent updates to the continuous slot state associated with each address. Under a restricted RAM-Net construction, we prove that soft address assignments extend hard quantized matching to a separable read-write overlap that locally approximates full-attention similarity and supports recurrent aggregation. These connections further enable Transformer-to-RAM-Net weight migration through a new path based on a soft-quantized intermediate construction. Across nine pretrained Transformer models from 0.3B to 7B parameters, RAM-Net recovers an average of 87.1% of the teachers' accuracy gains over random guessing across six commonsense and knowledge tasks using only a 500M-token budget per model.
Figures & tables
| Family | Teacher | Model | Tokens | Ctx. | Wiki PPL | ARC-E | ARC-C | PIQA | SciQ | COPA | Hella. | ||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| FLA | transformer-340M- 10B | Teacher | 10B | 4K | 30.05 | 57.4 | 24.1 | 65.6 | 82.5 | 73.0 | 32.1 | 55.8 | 86.5 |
| RAM-Net | 0.5B | 2K | 35.14 | 54.7 | 23.7 | 64.7 | 80.2 | 65.0 | 31.0 | 53.2 | |||
| transformer-1.3B- 100B | Teacher | 100B | 2K | 17.66 | 56.1 | 23.9 | 70.1 | 85.3 | 75.0 | 38.5 | 58.1 | 86.4 | |
| RAM-Net | 0.5B | 2K | 21.93 | 52.7 | 24.8 | 68.3 | 78.8 | 71.0 | 35.6 | 55.2 | |||
| transformer-2.7B- 100B | Teacher | 100B | 2K | 15.23 | 59.4 | 26.5 | 71.9 | 87.9 | 75.0 | 42.0 | 60.4 | 89.2 | |
| RAM-Net | 0.5B | 2K | 19.82 | 56.5 | 24.6 | 70.3 | 79.9 | 74.0 | 38.3 | 57.3 |
| Method | FLA-0.3B | FLA-1.3B | FLA-2.7B | Qwen-0.5B | Qwen-1.5B | Qwen-3B | Red.-3B | Mist.-7B | LLaMA-7B |
|---|---|---|---|---|---|---|---|---|---|
| r/w projection | 0.561 | 0.563 | 0.568 | 0.625 | 0.617 | 0.601 | 0.611 | 0.619 | 0.629 |
| QK projection adapter | 0.570 | 0.558 | 0.571 | 0.627 | 0.629 | 0.623 | 0.601 | 0.615 | 0.615 |
| Random codebook | 0.441 | 0.440 | 0.437 | 0.486 | 0.476 | 0.498 | 0.443 | 0.430 | 0.389 |
| Orth. Rand. codebook | 0.487 | 0.515 | 0.525 | 0.574 | 0.566 | 0.560 | 0.563 | 0.568 | 0.586 |
| Shared -means codebook | 0.415 | 0.412 | 0.419 | 0.466 | 0.453 | 0.481 | 0.421 | 0.411 | 0.377 |
| Per-factor PQ codebook | 0.412 | 0.423 | 0.431 | 0.487 | 0.476 | 0.487 | 0.446 | 0.445 | 0.407 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Family | Teacher | Layer | Hidden size | Head | KV Head | Attn. | Native ctx. | |
| FLA | transformer-340M-10B | 24 | 1024 | 32 | 32 | 32 | MHA | 4K |
| transformer-1.3B-100B | 24 | 2048 | 32 | 32 | 64 | MHA | 2K | |
| transformer-2.7B-100B | 32 | 2560 | 20 | 20 | 128 | MHA | 2K | |
| Qwen2.5 | 0.5B-Base | 24 | 896 | 14 | 2 | 64 | GQA | 32K |
| 1.5B-Base | 28 | 1536 | 12 | 2 | 128 | GQA | 32K | |
| 3B-Base | 36 | 2048 | 16 | 2 | 128 | GQA | 32K |
| Phase | Stage | Main Loss | Trainable Parameters | Tokens | Warmup | LR |
| Intermediate | Low-rank & positional alignment | Attention KL | QK adapters, log-scales | 7.209M | 0.360M | |
| Circular soft quantization | – | – | 0 | – | – | |
| Recovery | Head-level | Head-output NMSE | , , scales, | 49.152M | 1.638M | |
| Attn.-block local | Residual-update NMSE | Attention branch | 61.440M | 2.048M | ||
| Attn.-block stack | Residual-update NMSE | Attention branch | 36.864M | 1.638M | ||
| Full-model | Logit KL + CE | All parameters | 345.375M | 0.819M |
| Family | Teacher | Wiki PPL | ARC-E | ARC-C | PIQA | SciQ | COPA | Hella. | Avg. | R |
|---|---|---|---|---|---|---|---|---|---|---|
| FLA | transformer-340M-10B | 35.28 (+0.1) | 55.8 (+1.1) | 23.9 (+0.2) | 64.3 (-0.4) | 80.3 (+0.1) | 66.0 (+1.0) | 31.1 (+0.1) | 53.6 (+0.4) | 87.7 (+1.2) |
| transformer-1.3B-100B | 21.83 (-0.1) | 52.3 (-0.4) | 23.7 (-1.1) | 67.7 (-0.6) | 78.0 (-0.8) | 72.0 (+1.0) | 35.8 (+0.2) | 54.9 (-0.3) | 86.4 (-0.1) | |
| transformer-2.7B-100B | 20.09 (+0.3) | 56.8 (+0.3) | 25.4 (+0.8) | 69.6 (-0.7) | 79.4 (-0.5) | 76.0 (+2.0) | 37.8 (-0.5) | 57.5 (+0.2) | 89.6 (+0.4) | |
| Qwen2.5 | 0.5B-Base | 23.55 (+0.2) | 64.1 (+0.5) | 28.8 (+2.4) | 68.4 (-0.1) | 81.9 (0.0) | 69.0 (-4.0) | 36.8 (-0.1) | 58.2 (-0.2) | 86.1 (+6.9) |
| 1.5B-Base | 17.30 (+1.0) | 74.8 (-0.5) | 40.2 (-1.3) | 74.7 (+0.4) | 88.4 (-1.4) | 77.0 (-5.0) | 44.7 (-0.9) | 66.6 (-1.5) | 90.2 (-4.8) | |
| 3B-Base | 13.89 (+0.3) | 77.2 (0.0) | 44.2 (+0.5) | 77.0 (0.0) | 91.6 (-1.3) | 81.0 (+2.0) | 50.3 (-0.4) | 70.2 (+0.1) | 92.3 (+0.8) |
| Readout rule | FLA-0.3B | FLA-1.3B | FLA-2.7B | Qwen-0.5B | Qwen-1.5B | Qwen-3B | Red.-3B | Mist.-7B | LLaMA-7B |
|---|---|---|---|---|---|---|---|---|---|
| w/o | 0.277 | 0.286 | 0.308 | 0.292 | 0.293 | 0.312 | 0.312 | 0.283 | 0.253 |
| w/ | 0.259 | 0.273 | 0.300 | 0.290 | 0.297 | 0.311 | 0.319 | 0.280 | 0.257 |
| Wiki PPL | ARC-E | ARC-C | PIQA | SciQ | COPA | Hella. | Avg. | |
|---|---|---|---|---|---|---|---|---|
| w/o | 26.89 | 59.1 | 24.7 | 66.9 | 82.5 | 67.0 | 33.4 | 55.6 |
| w/ | 27.10 | 59.1 | 26.1 | 67.5 | 82.1 | 73.0 | 33.3 | 56.8 |
| Metric | Method | FLA-0.3B | FLA-1.3B | FLA-2.7B | Qwen-0.5B | Qwen-1.5B | Qwen-3B | Red.-3B | Mist.-7B | LLaMA-7B |
|---|---|---|---|---|---|---|---|---|---|---|
| LCC | Teacher | 17.0 | 48.3 | 50.7 | 10.0 | 14.7 | 14.1 | 44.7 | 70.0 | 66.7 |
| RAM-Net | 19.6 | 36.4 | 36.4 | 18.5 | 12.7 | 15.5 | 32.7 | 33.4 | 24.8 | |
| MultiNews | Teacher | 10.5 | 6.6 | 6.3 | 21.4 | 25.2 | 24.5 | 20.1 | 15.5 | 13.8 |
| RAM-Net | 9.6 | 7.2 | 8.4 | 12.9 | 13.9 | 15.9 | 8.6 | 12.1 | 11.5 | |
| SAMSum | Teacher | 6.6 | 6.2 | 4.9 | 37.3 | 42.4 | 45.5 | 0.0 | 28.2 | 31.4 |
| RAM-Net | 8.0 | 9.3 | 10.3 | 20.3 | 28.7 | 29.5 | 0.1 | 22.1 | 23.9 |
| Order | Low-rank KL | Quant. JS | Head NMSE | Logit KL | Wiki PPL | ARC-E | ARC-C | PIQA | SciQ | COPA | Hella. | Avg. | R |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 4 (256 slots) | 1.957 | 0.253 | 0.177 | 0.500 | 16.58 | 73.8 | 38.9 | 73.6 | 87.4 | 82.0 | 45.6 | 66.9 | 90.8 |
| 5 (1024 slots) | 1.748 | 0.293 | 0.197 | 0.465 | 16.33 | 75.3 | 41.5 | 74.3 | 89.8 | 82.0 | 45.6 | 68.1 | 95.0 |
| 6 (4096 slots) | 1.555 | 0.339 | 0.209 | 0.476 | 16.01 | 74.4 | 41.6 | 74.0 | 89.2 | 77.0 | 45.4 | 66.9 | 91.8 |
| Method | Checkpoint | Tokens (M) | Logit KL | Wiki | ARC-E | ARC-C | PIQA | SciQ | COPA | Hella. | Avg. | R |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Teacher | Original | – | – | 12.60 | 75.2 | 41.1 | 75.6 | 94.2 | 83.0 | 50.2 | 69.9 | – |
| Complete | Head-Rec. end | 56.4 | – | 77.56 | 47.1 | 25.1 | 64.2 | 69.9 | 69.0 | 35.2 | 51.7 | 43.8 |
| Block-Rec. end | 154.7 | – | 19.44 | 72.1 | 37.6 | 73.6 | 87.2 | 79.0 | 44.1 | 65.6 | 86.3 | |
| Final | 500.0 | 0.164 | 16.33 | 75.3 | 41.5 | 74.3 | 89.8 | 82.0 | 45.6 | 68.1 | 95.0 | |
| Full-Model only | Matched head | 56.4 | 0.419 | 23.90 | 65.1 | 31.5 | 70.2 | 83.0 | 77.0 | 40.7 | 61.2 | 71.1 |
| Matched block | 154.7 | 0.278 | 19.35 | 70.1 | 35.7 | 73.6 | 87.3 | 78.0 | 43.2 | 64.7 | 82.6 |