cs.LGOct 7, 2026

Dual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV Caches

Authors: Sunjoo Whang, Jungjun Oh, Minsung Kim, Dongho Seo, Jisu Shin, Gregory Kielian, Hoi-Jun Yoo, Sangjin Kim

Organizations: KAIST · GIST · Google Research

Abstract

Long inputs and extended generation increase the storage and access costs of the key-value (KV) cache. Low-bit quantization reduces storage and memory traffic, while query-channel pruning can further reduce key-cache reads. Rotation-based quantization redistributes the energy of key outliers across channels. To maintain computational invariance, the same orthogonal transform must be applied to queries, preserving query-key dot products. However, this rotation can disperse query energy, weakening the separation between a few large components to retain and many small ones to prune. We introduce Dual-QK, which uses paired non-orthogonal query and key transforms to address this conflict. Using calibrated query and key statistics, Dual-QK combines partial key whitening with a query-aligned basis to balance key scales for INT2 quantization and concentrate query energy for dynamic channel pruning. Channel-0 protection and bucket-relative RoPE support low-bit accuracy over long contexts. Experiments on four models across five generative benchmarks and long-context retrieval tasks show improved accuracy over OSCAR on most tasks at 40% query-channel sparsity. At a 128K context, Dual-QK provides 6.8×6.8\times KV-cache compression and an estimated 8.3×8.3\times reduction in KV read volume relative to unpruned BF16. Under the evaluated configurations, our SGLang implementation achieves up to 3.75×3.75\times the decoding throughput of unpruned BF16.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Read What Matters: Query-Adaptive Quantization for KV Caches

    Oct 8, 2026Siddharth Bhandari, Lucas Gretta, Krishna Balasubramanian +1KV-Cache QuantizationKV-Cache Compression

  2. KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation

    Sep 21, 2026Sihyeon Ha, Jaeho Lee, Yo-Seb JeonLow-Rank CompressionKV-Cache Compression

  3. Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms

    Aug 4, 2026Samuel Fernández-Menduiña, Amir Ziashahabi, Eduardo Pavez +2KV CachingVector Quantization