Long inputs and extended generation increase the storage and access costs of the key-value (KV) cache. Low-bit quantization reduces storage and memory traffic, while query-channel pruning can further reduce key-cache reads. Rotation-based quantization redistributes the energy of key outliers across channels. To maintain computational invariance, the same orthogonal transform must be applied to queries, preserving query-key dot products. However, this rotation can disperse query energy, weakening the separation between a few large components to retain and many small ones to prune. We introduce Dual-QK, which uses paired non-orthogonal query and key transforms to address this conflict. Using calibrated query and key statistics, Dual-QK combines partial key whitening with a query-aligned basis to balance key scales for INT2 quantization and concentrate query energy for dynamic channel pruning. Channel-0 protection and bucket-relative RoPE support low-bit accuracy over long contexts. Experiments on four models across five generative benchmarks and long-context retrieval tasks show improved accuracy over OSCAR on most tasks at 40% query-channel sparsity. At a 128K context, Dual-QK provides 6.8× KV-cache compression and an estimated 8.3× reduction in KV read volume relative to unpruned BF16. Under the evaluated configurations, our SGLang implementation achieves up to 3.75× the decoding throughput of unpruned BF16.
Figures & tables
Figure 1: Conceptual comparison at the same bit width and retained-channel budget. (a) Q-PCA concentrates query energy, whereas (b) Hadamard rotation mitigates key outliers. (c) Dual-QK targets both properties using reciprocal Q/K transforms. Hatched channels are omitted from the dot product.
Figure 2: Query (top) and key (bottom) magnitudes from Llama-3.1-8B-Instruct (layer 29, KV group 6, query head 26) in the original, Q-PCA, Hadamard, and Dual-QK coordinates. Q-PCA concentrates queries but leaves key imbalance. Hadamard spreads both, whereas Dual-QK concentrates queries while reducing key imbalance. Magnitude axes are scaled independently across panels.
Model
Method
KV bits per element
Benchmark score (%)
Capacity
Bandwidth
GPQA-Diamond
HumanEval
LCB v6
MATH-500
AIME25
Llama-3.1-8B -Instruct
BF16
16
16
$$26.2\pm{\scriptstyle1.5}
$$66.7\pm{\scriptstyle2.5}
$$16.5\pm{\scriptstyle1.4}
$$49.4\pm{\scriptstyle1.5}
$$1.3\pm{\scriptstyle1.8}
TurboQuant
3.13
2.53
$$17.7\pm{\scriptstyle2.4}
$$24.3\pm{\scriptstyle1.4}
$$3.8\pm{\scriptstyle0.5}
$$13.8\pm{\scriptstyle1.1}
$$0.7\pm{\scriptstyle1.5}
KIVI
3.01
2.41
$$19.2\pm{\scriptstyle2.7}
$$55.2\pm{\scriptstyle4.3}
$$9.8\pm{\scriptstyle2.6}
$$28.9\pm{\scriptstyle1.4}
$$0.7\pm{\scriptstyle1.5}
OSCAR
2.28
1.88
$$16.4\pm{\scriptstyle3.0}
$$28.2\pm{\scriptstyle2.2}
$$3.2\pm{\scriptstyle1.0}
$$16.4\pm{\scriptstyle0.4}
\mathbf{1.3}\pm{\scriptstyle\mathbf{1.8}}
Dual-QK (ours)
2.34
1.93
\mathbf{23.2}\pm{\scriptstyle\mathbf{1.9}}
\mathbf{64.3}\pm{\scriptstyle\mathbf{2.9}}
\mathbf{14.2}\pm{\scriptstyle\mathbf{1.3}}
\mathbf{43.0}\pm{\scriptstyle\mathbf{2.0}}
$$0.7\pm{\scriptstyle1.5}
Table 1: Generative accuracy (%) with 40% query-channel sparsity for quantized methods. BF16 is unpruned. Scores report mean ± standard deviation over five seeds (three for Ministral-3). Capacity and bandwidth are effective bits per KV element at 128K context. Bold marks the highest mean among quantized methods within each model.
Model
Method
KV bits per element
Context length
Capacity
Bandwidth
4K
8K
16K
32K
64K
128K
Llama-3.1- 8B-Instruct
BF16
16
16
$$99.958\pm{\scriptstyle0.072}
$$99.875\pm{\scriptstyle0.000}
$$99.958\pm{\scriptstyle0.072}
$$99.958\pm{\scriptstyle0.072}
$$99.292\pm{\scriptstyle0.260}
$$94.833\pm{\scriptstyle0.402}
OSCAR
2.28
1.88
$$65.042\pm{\scriptstyle1.018}
$$53.583\pm{\scriptstyle1.563}
$$51.250\pm{\scriptstyle0.433}
$$42.875\pm{\scriptstyle1.682}
$$39.542\pm{\scriptstyle1.751}
$$25.208\pm{\scriptstyle0.591}
Dual-QK (ours)
2.34
1.93
\mathbf{92.375}\pm{\scriptstyle\mathbf{0.375}}
\mathbf{92.500}\pm{\scriptstyle\mathbf{0.760}}
\mathbf{89.000}\pm{\scriptstyle\mathbf{0.217}}
\mathbf{83.667}\pm{\scriptstyle\mathbf{0.688}}
\mathbf{78.917}\pm{\scriptstyle\mathbf{0.641}}
\mathbf{58.875}\pm{\scriptstyle\mathbf{1.474}}
Qwen3-4B -Thinking -2507
BF16
16
16
$$100.000\pm{\scriptstyle0.000}
$$98.958\pm{\scriptstyle0.144}
$$98.333\pm{\scriptstyle0.402}
$$96.833\pm{\scriptstyle0.402}
$$96.708\pm{\scriptstyle0.564}
$$94.000\pm{\scriptstyle0.125}
OSCAR
2.28
1.88
$$78.375\pm{\scriptstyle0.820}
$$72.000\pm{\scriptstyle0.375}
$$61.167\pm{\scriptstyle0.439}
$$54.958\pm{\scriptstyle0.832}
$$32.708\pm{\scriptstyle1.583}
$$2.000\pm{\scriptstyle0.000}
Table 2: RULER NIAH accuracy (%) across context lengths with 40% query-channel sparsity for quantized methods. BF16 is the unpruned reference. Entries report mean ± sample standard deviation across three NIAH data draws, each containing 25 examples per subtask across eight subtasks. Capacity and bandwidth are effective bits per KV element at a 128K context. Bold denotes the highest mean among quantized methods.
Figure 3: Accuracy on GPQA-Diamond and HumanEval as query-channel pruning increases for Qwen3-8B and Qwen3-4B-Thinking-2507. Dashed lines indicate the unpruned BF16 references. Solid curves show OSCAR and Dual-QK with quantized KV caches, including at zero pruning. Error bars indicate standard deviation over three seeds.
Configuration
GPQA-D
HumanEval
LCB v6
MATH-500
AIME25
BF16 reference
$$56.1\pm{\scriptstyle2.2}
$$90.1\pm{\scriptstyle1.3}
$$50.2\pm{\scriptstyle2.9}
$$96.4\pm{\scriptstyle0.3}
$$65.3\pm{\scriptstyle3.0}
Full Dual-QK ( α=0.5 )
$$55.7\pm{\scriptstyle1.9}
$$90.6\pm{\scriptstyle2.0}
$$36.2\pm{\scriptstyle1.2}
$$95.0\pm{\scriptstyle0.3}
$$53.3\pm{\scriptstyle2.4}
w/o channel-0 protection
$$46.5\pm{\scriptstyle4.5}
$$71.5\pm{\scriptstyle1.9}
$$16.5\pm{\scriptstyle1.2}
$$87.3\pm{\scriptstyle1.0}
$$24.4\pm{\scriptstyle1.9}
w/o bucket-relative RoPE
$$33.2\pm{\scriptstyle0.3}
$$51.4\pm{\scriptstyle2.5}
$$8.4\pm{\scriptstyle0.8}
$$51.1\pm{\scriptstyle1.1}
$$2.2\pm{\scriptstyle1.9}
w/o diagonal refinement ( D=I )
$$52.5\pm{\scriptstyle4.8}
$$87.6\pm{\scriptstyle2.8}
$$35.4\pm{\scriptstyle1.9}
$$89.5\pm{\scriptstyle5.4}
$$51.1\pm{\scriptstyle9.6}
Q-PCA base ( α=0 )
$$45.3\pm{\scriptstyle3.8}
$$75.2\pm{\scriptstyle3.9}
$$15.3\pm{\scriptstyle1.3}
$$78.7\pm{\scriptstyle11.6}
$$17.8\pm{\scriptstyle1.9}
Table 3: Component ablation of Dual-QK on Qwen3-8B with INT2 KV quantization and 40% query-channel sparsity. BF16 is unpruned. Entries report mean ± standard deviation. BF16 and Full Dual-QK reuse the five-seed results from Table 1 . The remaining configurations use three seeds.
Model
Method
GPQA-D
HumanEval
LCB v6
MATH-500
AIME25
Mean
Llama-3.1- 8B-Instruct
BF16
$$26.2\pm{\scriptstyle1.5}
$$66.7\pm{\scriptstyle2.5}
$$16.5\pm{\scriptstyle1.4}
$$49.4\pm{\scriptstyle1.5}
$$1.3\pm{\scriptstyle1.8}
32.0
TurboQuant
\mathbf{26.8}\pm{\scriptstyle\mathbf{3.1}}
$$59.3\pm{\scriptstyle2.9}
$$12.5\pm{\scriptstyle1.3}
$$39.9\pm{\scriptstyle1.1}
$$0.0\pm{\scriptstyle0.0}
27.7
OSCAR
$$24.3\pm{\scriptstyle1.7}
$$65.6\pm{\scriptstyle3.0}
$$15.3\pm{\scriptstyle0.9}
$$50.8\pm{\scriptstyle0.5}
\mathbf{2.0}\pm{\scriptstyle\mathbf{1.8}}
31.6
Dual-QK (ours)
$$26.4\pm{\scriptstyle1.9}
\mathbf{66.0}\pm{\scriptstyle\mathbf{1.7}}
\mathbf{16.5}\pm{\scriptstyle\mathbf{0.7}}
\mathbf{50.9}\pm{\scriptstyle\mathbf{0.6}}
$$1.3\pm{\scriptstyle3.0}
32.2
Qwen3-8B
BF16
$$56.1\pm{\scriptstyle2.2}
$$90.1\pm{\scriptstyle1.3}
$$50.2\pm{\scriptstyle2.9}
$$96.4\pm{\scriptstyle0.3}
$$65.3\pm{\scriptstyle3.0}
71.6
TurboQuant
$$47.1\pm{\scriptstyle2.2}
$$75.7\pm{\scriptstyle1.6}
$$24.7\pm{\scriptstyle1.7}
$$92.4\pm{\scriptstyle0.6}
$$50.7\pm{\scriptstyle2.8}
58.1
Table 4: Generative accuracy (%) without query-channel pruning . Benchmark entries report mean ± standard deviation over five seeds. BF16 is the unquantized reference. Bold denotes the highest score among quantized methods within each model.
Method
Sparsity
Calibration
GPQA-D
HumanEval
LCB v6
MATH-500
AIME25
BF16
—
—
$$63.0\pm{\scriptstyle2.4}
$$95.5\pm{\scriptstyle0.9}
$$51.7\pm{\scriptstyle0.4}
$$97.5\pm{\scriptstyle0.3}
$$72.2\pm{\scriptstyle3.8}
Dual-QK
0%
WikiText
\mathbf{62.3}\pm{\scriptstyle\mathbf{1.6}}
\mathbf{94.9}\pm{\scriptstyle\mathbf{0.4}}
\mathbf{44.5}\pm{\scriptstyle\mathbf{0.9}}
\mathbf{96.1}\pm{\scriptstyle\mathbf{0.2}}
\mathbf{47.8}\pm{\scriptstyle\mathbf{1.9}}
( ours )
MMLU
$$61.3\pm{\scriptstyle2.8}
$$94.7\pm{\scriptstyle0.7}
$$45.3\pm{\scriptstyle2.9}
$$95.7\pm{\scriptstyle0.4}
$$57.8\pm{\scriptstyle1.9}
GPQA-D
$$62.3\pm{\scriptstyle1.8}
$$94.9\pm{\scriptstyle0.9}
$$45.0\pm{\scriptstyle1.5}
$$96.7\pm{\scriptstyle0.8}
$$60.0\pm{\scriptstyle5.8}
40%
WikiText
\mathbf{58.1}\pm{\scriptstyle\mathbf{1.3}}
\mathbf{92.9}\pm{\scriptstyle\mathbf{1.8}}
\mathbf{38.9}\pm{\scriptstyle\mathbf{2.3}}
\mathbf{93.9}\pm{\scriptstyle\mathbf{0.8}}
\mathbf{51.1}\pm{\scriptstyle\mathbf{1.9}}
MMLU
$$56.6\pm{\scriptstyle1.5}
$$94.7\pm{\scriptstyle0.4}
$$39.4\pm{\scriptstyle1.6}
$$93.7\pm{\scriptstyle1.0}
$$50.0\pm{\scriptstyle0.0}
Table 5: Calibration-domain ablation of Dual-QK on Qwen3-4B-Thinking-2507. Results are reported as mean ± standard deviation over three seeds. Sparsity denotes query-channel sparsity. BF16 is the unpruned reference. Bold highlights WikiText, our default calibration dataset, and its results.
Figure 5: Channel profiles across four models. (a ∼ d) Base key energy at different α . (e ∼ h) Query energy in the original, Q-PCA, and deployed dual coordinates. (i ∼ l) Channel-wise shares of key reconstruction error (dotted) and query-weighted error terms (solid), before channel-0 protection, including a Hadamard control. The first two rows use calibration moments and the last uses WikiText-2 test activations. Error shares are pooled across layers and KV heads.
Model
Method
KV bits per element
Benchmark score (%)
Capacity
Bandwidth
GPQA-Diamond
HumanEval
LCB v6
MATH-500
AIME25
Llama-3.1-8B -Instruct
BF16
16
16
$$26.2\pm{\scriptstyle1.5}
$$66.7\pm{\scriptstyle2.5}
$$16.5\pm{\scriptstyle1.4}
$$49.4\pm{\scriptstyle1.5}
$$1.3\pm{\scriptstyle1.8}
TurboQuant
3.13
3.13
\mathbf{26.8}\pm{\scriptstyle\mathbf{3.1}}
$$59.3\pm{\scriptstyle2.9}
$$12.5\pm{\scriptstyle1.3}
$$39.9\pm{\scriptstyle1.1}
$$0.0\pm{\scriptstyle0.0}
KIVI-2
3.01
3.01
$$23.8\pm{\scriptstyle2.8}
$$61.7\pm{\scriptstyle1.7}
$$14.8\pm{\scriptstyle1.8}
$$43.6\pm{\scriptstyle1.0}
$$0.0\pm{\scriptstyle0.0}
OSCAR
2.28
2.28
$$24.3\pm{\scriptstyle1.7}
$$65.6\pm{\scriptstyle3.0}
$$15.3\pm{\scriptstyle0.9}
$$50.8\pm{\scriptstyle0.5}
\mathbf{2.0}\pm{\scriptstyle\mathbf{1.8}}
Dual-QK (ours)
2.34
2.34
$$26.4\pm{\scriptstyle1.9}
\mathbf{66.0}\pm{\scriptstyle\mathbf{1.7}}
\mathbf{16.5}\pm{\scriptstyle\mathbf{0.7}}
\mathbf{50.9}\pm{\scriptstyle\mathbf{0.6}}
$$1.3\pm{\scriptstyle3.0}
Appendix
Table 9: Generative accuracy (%) without query-channel pruning. Scores report mean ± standard deviation over five seeds. Capacity and bandwidth are effective bits per KV element at 128K context. Bold marks the highest mean among quantized methods within each model.
Model
Method
Qasper
QMSum
Multi News
TREC
Trivia QA
SAMSum
LCC
Repo Bench-P
Mean
Llama-3.1-8B -Instruct
BF16
39.59
23.79
24.87
56.00
80.44
33.84
35.94
30.01
40.56
TurboQuant
14.30
18.66
12.99
46.50
54.04
20.86
18.36
17.99
25.46
OSCAR
18.03
18.29
15.45
49.50
57.49
24.26
18.13
19.62
27.60
Dual-QK (ours)
38.97
22.99
24.31
54.00
75.00
33.60
26.95
25.65
37.68
Qwen3-8B
BF16
46.72
23.80
23.86
58.63
89.13
41.04
25.98
23.51
41.58
TurboQuant
2.50
6.72
11.74
23.50
0.04
1.23
14.48
12.69
9.11
Appendix
Table 10: LongBench scores with 40% query-channel sparsity for quantized methods. BF16 is unpruned. Mean averages the eight task scores. Bold marks the highest score among quantized methods within each model.
Model
Method
Qasper
QMSum
Multi News
TREC
Trivia QA
SAMSum
LCC
Repo Bench-P
Mean
Llama-3.1-8B -Instruct
BF16
39.59
23.79
24.87
56.00
80.44
33.84
35.94
30.01
40.56
TurboQuant
39.21
22.60
22.94
55.25
75.95
30.26
27.67
25.68
37.44
OSCAR
36.91
22.67
23.99
53.50
76.53
34.10
35.33
28.31
38.92
Dual-QK (ours)
36.06
22.65
24.67
56.50
77.32
33.87
34.28
28.38
39.22
Qwen3-8B
BF16
46.72
23.80
23.86
58.63
89.13
41.04
25.98
23.51
41.58
TurboQuant
37.08
20.96
23.99
66.50
73.78
36.17
20.99
19.88
37.42
Appendix
Table 11: LongBench scores without query-channel pruning. Mean averages the eight task scores. Bold marks the highest score among quantized methods within each model.
What limits KV-cache compression at extreme bit-rates? We argue that it is not the choice of compression scheme, but how its budget is allocated across attention heads. Existing methods apply rank and bit-width uniformly, ignoring that each head has a different optimal mix of rank truncation and quantization. We show that co-optimizing rank and bit-width per head, using only standard low-rank projection and scalar quantization, dominates uniform allocation, with the largest gains at low bit-rates. Our method, KV-COBRA (Co-Optimized Bit-Rank Allocation), formalizes this as a resource-allocation problem: it balances rank-truncation loss against quantization loss within each head, then redistributes budget across heads to minimize total distortion. A fused Hadamard rotation equalizes per-channel variance, and reordering the SVD basis by attention-KL importance makes the solver query-aware. The same allocator extends to joint K+V compression. On perplexity, zero-shot, and long-context benchmarks from 0.5 to 4 bits per dimension (bpd), KV-COBRA shows the smallest accuracy degradation among evaluated methods at low bpd, with no per-token overhead.
Sihyeon Ha, Jaeho Lee, Yo-Seb Jeon
Pohang University of Science and Technology (POSTECH)
Long-context LLM decoding reads the key-value (KV) cache at every step. Loading it takes longer than computing attention over it, so throughput is bandwidth-bound. Hence, reducing the cache size can raise both decoding speed and serving capacity. The challenge is to reduce cache size while preserving the attention products, keeping reconstruction cheap, and using a fixed per-token bit count. At two bits per element, the most competitive methods rely on orthogonal transforms. However, existing techniques are either data-oblivious or use the query statistics without deriving the transform from a distortion criterion. Moreover, they rely on transforms built on top of random or Hadamard rotations, which equalize variances across entries rather than compacting energy, and fixed-width scalar quantizers, which are suboptimal at low rates. In this paper, we formulate KV cache quantization as a transform coding problem in which distortion is the error in the attention products. We derive closed-form optimal transforms for keys and values from calibration statistics, under a high-resolution model. We show that the optimal key transform is not orthogonal and satisfies a generalized Parseval relation: the attention-aware distortion becomes mean-squared error (MSE) in the transform domain. Thus, we can use MSE-optimal vector quantizers applied directly to the transformed key coefficients. To meet the fixed-width layout requirement, we show that grouping coefficients into equal-volume partitions makes equal-size codebooks attain the variable-rate optimum under the same high-resolution model. At two bits per element, our method, termed NOVA-KV, recovers most of the long-context retrieval accuracy lost by scalar quantization methods at comparable throughput.
Samuel Fernández-Menduiña, Amir Ziashahabi, Eduardo Pavez +2
Department of Electrical and Computer Engineering, University of Southern California
Existing low-bit KV-cache quantizers often treat each cached key as a flat vector. Under RoPE, however, a key's contribution to a future attention logit decomposes into a position-dependent sum over two-dimensional frequency blocks. This makes key-cache quantization a block-wise bit-allocation problem: high-energy RoPE blocks are more sensitive to quantization error and should receive more bits. We introduce Block-GTQ, a RoPE-aware bit allocator for key-cache quantization built on TurboQuant-MSE(TQ-MSE). For each layer and KV head, Block-GTQ computes a label-free energy score for each RoPE block and greedily allocates integer bit widths by marginal gain. Under matched K/V bit budgets, Block-GTQ better preserves RoPE query-key logits on a ten-model diagnostic panel, cutting per-layer MAE by 32-80% at 2 and 3 b/dim K-only quantization and winning all 367/367 layer comparisons against uniform TQ-MSE. These fidelity gains translate to stronger downstream long-context retrieval, understanding, and reasoning. At K2V2 on Llama-3.1-8B-Instruct, Block-GTQ raises the six-task NIAH average from 70.6 to 97.4, and the LongBench-EN average from 36.87 to 53.31. On AIME 2024/2025 with DeepSeek-R1-Distill-Qwen-7B, without an fp16 recent-key buffer, Block-GTQ at K3V2 scores 51.7/37.5, close to fp16's 54.2/37.9, whereas uniform TQ-MSE collapses to 0.0/0.0. We further implement a packed-cache serving path. On a single H800 GPU with Qwen2.5-3B-Instruct, packed K3V3 achieves 3.24x KV-cache compression with fp16-comparable quality, runs 1.34x faster than fp16 FlashAttention2 at 128K context, reduces peak memory from 56.31 GB to 19.85 GB, and remains feasible at 256K and 512K where fp16 OOMs. Code is available at https://github.com/JIA-Lab-research/blockgtq.
Fengfeng Liang, Yuechen Zhang, Jiaya Jia
1Hong Kong University of Science and Technology · 2The Chinese University of Hong Kong · 3MiMo, Xiaomi Corporation