Read What Matters: Query-Adaptive Quantization for KV Caches
Organizations: Amazon · University of California, Berkeley
Abstract
KV-cache entries are stored before their future queries are known, but each decoding query needs precision in different places. We study this mismatch using separate budgets for retained bits and bits fetched per query. ReadKV stores each key and value in a progressive code whose prefixes support different reconstruction precisions. For each query, it allocates key-channel prefixes using the query, computes attention from the reconstructed keys, and then allocates value-token prefixes using that attention. Stored entries remain unchanged. Each stage optimizes a calibrated distortion objective under a fixed budget; we prove exact allocation under diminishing refinement gains and relate these objectives to attention-output error. We also exhibit a finite-dimensional attention family where query-dependent access strictly outperforms every query-independent reader at the same read budget, even with unrestricted competing encoders and decoders. Across six base models, reading four bits on average from an eight-bit cache increases C4 perplexity by at most 0.66%, using about one quarter of the logical reads and half the retained capacity of a 16-bit cache. It is consistently more accurate than storing and fully reading four bits at the same payload-read budget. Retaining more bits than each query fetches is aimed at long-context decoding, where the cache bytes moved per step, rather than the weights, dominate cost. Long-context question answering and retrieval on two instruction-tuned models provide additional quality evidence. On the tested 8K-token, batch-one, single-layer workload on an NVIDIA A10G, a restricted eight-bit ReadKV reader with a two-bit mean payload-read budget has 39% lower latency than the tested TurboQuant codec.
Figures & tables
| Reader Class | Uses Query? | Use of Returned Bits in Address Selection |
|---|---|---|
| Query-oblivious fixed pattern | No | None; every address is fixed in advance |
| Query-independent branching | No | Each new address may use all previously returned bits |
| : one batch | Yes | None; the complete batch is selected before reading |
| , : at most batches | Yes | A later batch may use bits returned by earlier batches |
| Fully branching | Yes | Each new address may use all previously returned bits |
| Each method cell reports (%) on top and logical reads / retained KV capacity (% of Dense ) below. | |||
| Panel A: Qwen2.5 Models | |||
| Method | Qwen2.5 3B | Qwen2.5 7B | Qwen2.5 14B |
| Dense PPL | 11.36 | 10.13 | 8.74 |
| Full Reader | +1.59 [1.5pt] 25.6 / 25.1 | +1.34 [1.5pt] 25.4 / 25.1 | +1.68 [1.5pt] 25.3 / 25.1 |
| KIVI (2-bit K/V) [1pt] Liu et al.,2024 | +4.84 [1.5pt] 26.2 / 23.8 | +2.79 [1.5pt] 26.2 / 23.8 | +2.21 [1.5pt] 26.2 / 23.8 |
| SparQ [1pt] Ribar et al.,2024 | +1.27 [1.5pt] 24.7 / 100.0 | +1.44 [1.5pt] 24.7 / 100.0 | +0.81 [1.5pt] 24.7 / 100.0 |
| Each method cell reports (points) on top and logical reads / retained KV capacity (% of Dense ) below. | |||
| Panel A: Qwen2.5-7B-Instruct | |||
| Method | Qasper | HotpotQA | 2Wiki MultihopQA |
| Dense F1 | 41.13 | 48.34 | 45.30 |
| KIVI (2-bit K/V) [1pt] Liu et al.,2024 | +1.56 [1.5pt] 20.4 / 19.9 | +0.08 [1.5pt] 20.0 / 19.9 | +2.41 [1.5pt] 20.1 / 19.9 |
| TurboQuant K4/V4 † [1pt] Zandieh et al.,2026 | -0.06 [1.5pt] 36.7 / 36.7 | +0.42 [1.5pt] 36.7 / 36.7 | -0.55 [1.5pt] 36.7 / 36.7 |
| SparQ [1pt] Ribar et al.,2024 | +1.40 [1.5pt] 15.3 / 100.0 | -0.44 [1.5pt] 14.2 / 100.0 | +0.69 [1.5pt] 14.5 / 100.0 |
| Method | KV payload (MiB) | Latency (ms) | DRAM reads (MiB) | Total traffic (MiB) |
|---|---|---|---|---|
| Dense | 16.00 | 0.051 | 17.73 / 17.86 | 20.85 / 21.06 |
| TurboQuant K4/V4 ( Zandieh et al., 2026 ) | 4.19 | 0.154 | 7.28 / 7.92 | 9.38 / 10.84 |
| ReadKV | 8.00 | 0.093 | 2.80 / 2.74 | 7.47 / 7.47 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Domain or shape | Meaning |
|---|---|---|
| Partial-read model | ||
| ; | Dimension and the stored and query vectors in the worst-case inner-product problem. | |
| ; | , | Normalized storage and read rates, and their integer bit budgets. |
| Stored binary record produced before the query is known. | ||
| Public state | Code and memory-layout description fixed independently of the stored vector and query. | |
| Random variables | Public, encoder-local, and reader-local randomness, all independent of . | |
| Method | Qwen2.5 3B | Qwen2.5 7B | Qwen2.5 14B | Yi-1.5 6B | DeepSeek LLM-7B | Mistral 7B-v0.3 |
|---|---|---|---|---|---|---|
| Dense PPL | 11.36 | 10.13 | 8.74 | 8.96 | 8.49 | 7.62 |
| Top: (%) | ||||||
| Bottom: logical reads / retained KV capacity (% of Dense ) | ||||||
| ReadKV | +3.05 [1.5pt]22.7 / 55.6 | +1.93 [1.5pt]25.3 / 57.2 | +2.37 [1.5pt]20.1 / 54.3 | +7.82 [1.5pt]23.8 / 56.3 | +4.39 [1.5pt]24.4 / 56.8 | +8.39 [1.5pt]23.7 / 56.3 |
| KIVI (4-bit K/V) [1pt] Liu et al.,2024 | +0.07 [1.5pt]37.6 / 35.5 | +0.09 [1.5pt]37.6 / 35.5 | +0.06 [1.5pt]37.6 / 35.5 | +0.07 [1.5pt]37.6 / 35.5 | +0.10 [1.5pt]37.6 / 35.5 | +0.09 [1.5pt]37.6 / 35.5 |
| Lloyd-2 (norm-corrected) [1pt] Lloyd,1982 | +40.99 [1.5pt]13.1 / 13.0 | +33.83 [1.5pt]13.1 / 13.0 | not run | +159.37 [1.5pt]13.1 / 13.0 | +3766.75 [1.5pt]13.1 / 13.0 | +1366.07 [1.5pt]13.1 / 13.0 |
| Qwen2.5-7B-Instruct | Mistral-7B-Instruct-v0.3 | |||||
|---|---|---|---|---|---|---|
| Method | Qasper | HotpotQA | 2Wiki MultihopQA | Qasper | HotpotQA | 2Wiki MultihopQA |
| Dense F1 | 35.51 | 42.49 | 44.87 | 28.10 | 41.42 | 37.46 |
| Top: (points) Color : quality–read frontier | ||||||
| Bottom: logical reads / retained KV capacity (% of Dense ) | ||||||
| KIVI (2-bit K/V) [1pt] Liu et al.,2024 | +0.51 [1.5pt]21.2 / 20.7 | +2.59 [1.5pt]21.2 / 20.7 | +0.90 [1.5pt]21.2 / 20.7 | -0.39 [1.5pt]21.0 / 20.7 | -0.17 [1.5pt]20.9 / 20.7 | +0.32 [1.5pt]20.9 / 20.7 |
| TurboQuant K4/V4 † [1pt] Zandieh et al.,2026 | +1.14 [1.5pt]36.7 / 36.7 | +1.81 [1.5pt]36.7 / 36.7 | +1.00 [1.5pt]36.7 / 36.7 | +1.01 [1.5pt]35.4 / 35.4 | +0.17 [1.5pt]35.4 / 35.4 | +0.87 [1.5pt]35.4 / 35.4 |
| Qwen2.5-7B-Instruct | Mistral-7B-Instruct-v0.3 | |||
|---|---|---|---|---|
| Method | 4K | 8K | 4K | 8K |
| Dense recovered | 200 | 200 | 200 | 200 |
| Top: recovered keys (out of 200) Color : quality–read frontier | ||||
| Bottom: logical reads / retained KV capacity (% of Dense ) | ||||
| Full Reader | 200 [1.5pt]25.1 / 25.1 | 200 [1.5pt]25.1 / 25.0 | 200 [1.5pt]25.1 / 25.0 | 200 [1.5pt]25.0 / 25.0 |
| ReadKV | 200 [1.5pt]12.6 / 25.1 | 200 [1.5pt]12.6 / 25.0 | 200 [1.5pt]12.6 / 25.0 | 200 [1.5pt]12.5 / 25.0 |
| Allocation | Qwen3B | Qwen7B | Yi6B | DeepSeek7B |
|---|---|---|---|---|
| Uniform depth 2 | 21.66 | 19.91 | 39.34 | 568.35 |
| Calibration-only | 19.11 | 14.67 | 16.31 | 1 095.1 |
| Current query | 11.80 | 10.44 | 9.87 | 9.03 |
| Model | Equal K/V | Per head | Min. 1 | Paired | Total | |
|---|---|---|---|---|---|---|
| Qwen3B | 2 | +0.0729 | +0.0050 | +0.0012 | +0.0004 | +0.0795 |
| Qwen7B | 2 | +0.0575 | +0.0043 | +0.0061 | +0.0014 | +0.0694 |
| Qwen3B | 4 | +0.0032 | +0.0002 | +0.0004 | -0.0003 | +0.0035 |
| Qwen7B | 4 | +0.0034 | -0.0005 | +0.0002 | -0.0003 | +0.0028 |
| 2K tokens | 8K tokens | 32K tokens | ||||
|---|---|---|---|---|---|---|
| Method | ||||||
| Dense | 0.020 | 0.048 | 0.051 | 0.146 | 0.168 | 0.585 |
| TurboQuant K4/V4 codec [-1pt] Zandieh et al.,2026 | 0.055 | 0.143 | 0.154 | 0.518 | 0.625 | 1.990 |
| Full Reader | 0.038 | 0.076 | 0.070 | 0.171 | 0.265 | 0.727 |
| ReadKV | 0.063 | 0.095 | 0.092 | 0.179 | 0.287 | 0.719 |
| ReadKV | 0.066 | 0.096 | 0.093 | 0.183 | 0.304 | 0.740 |
| Method | DRAM reads | DRAM writes | Total |
|---|---|---|---|
| Dense | 17.73 / 17.86 | 3.12 / 3.19 | 20.85 / 21.06 |
| ReadKV | 2.72 / 2.71 | 4.63 / 4.71 | 7.35 / 7.42 |
| Full Reader | 4.70 / 4.66 | 4.19 / 4.25 | 8.90 / 8.91 |
| ReadKV , scheduled | 4.81 / 4.85 | 4.64 / 4.73 | 9.44 / 9.58 |
| ReadKV | 2.80 / 2.74 | 4.67 / 4.73 | 7.47 / 7.47 |
| ReadKV | 4.99 / 4.98 | 4.72 / 4.75 | 9.71 / 9.74 |