cs.CLOct 5, 2026

Shared Stopping Decisions Change Answers in HQQ Cache Quantization

Authors: Seunghui Jwa, Minsu Oh, Chanjun Park, Yeo-Chan Yoon

Organizations: Department of Artificial Intelligence, Jeju National University · School of Software, Soongsil University

Abstract

Language-model systems batch questions for throughput, but unrelated questions should not change a target's answer when its input and numerical execution are fixed. We study compression of the key and value cache, which stores attention representations reused during generation. With request-local groups, Transformers' Half-Quadratic Quantization (HQQ) backend updates compression parameters separately but uses a shared average error to decide when all updates stop. Replacing only the question batched with the target changes four-bit HQQ answers in 170/384 test comparisons across two models. Replaying the other execution's update counts reproduces its complete answer and cache fingerprints in every changed pair, in both directions. Computing the stopping mean in FP32 reduces cache differences but leaves answer changes. Native HQQ also changes confirmed numerical correctness in eight arithmetic pairs. Fixed iterations and request-local stopping remove observed companion dependence under matched controls. Request-local stopping remains sensitive to synthetic padding changes at the tensor level. Fixing the original iteration budget removes this decision path without tuning. Neither repair has an established quality advantage, and natural rebatching still changes answers. Request-independence audits must cover stopping decisions as well as quantization groups.

Figures & tables

Appendix figures & tables20 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

May 26, 2026cs.LG

Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression

We propose \textbf{Hurwitz Quaternion Multiplicative Quantization (HQMQ)}, a \textbf{calibration-free} method for KV cache compression of large language models. HQMQ treats each 4-element chunk of K or V as a quaternion and quantizes its unit direction to the \emph{product} qp⋅qsq_p \cdot q_s, where qpq_p ranges over the 24-element Hurwitz group 2T2T (the 24 vertices of the 24-cell on S3S^3, pairwise angle 60∘60^\circ) and qsq_s ranges over a per-(layer, head) secondary codebook of SS \emph{random} unit quaternions. The multiplicative composition yields 24S24S effective codewords at SS stored parameters; random initialization suffices because left-multiplication is an S3S^3 isometry, so seeded codebooks vary in end-task ppl by <1.5%<1.5\%. A per-batch median-multiplier outlier extraction step (C=3C{=}3, no calibration) handles modern outlier-heavy architectures. We evaluate on five modern open models: Mistral-7B (dense MHA), Llama-3-8B and Qwen2.5-7B and Qwen3-8B (dense GQA), and gpt-oss-20b (sparse MoE). On Mistral-7B and Qwen3-8B, HQMQ matches fp16 within 0.020.02--0.030.03 ppl points at ∼\sim5 bits. On Qwen2.5-7B and Qwen3-8B, where naive int4 collapses to 104+10^4{+} ppl, HQMQ + Med3×\times recovers fp16 quality within 0.020.02--0.100.10 ppl points at ∼\sim5 bits. HQMQ Pareto-dominates naive int by 33--1900×1900\times at matched bits across all five models, and downstream zero-shot accuracy matches fp16 at 3.793.79 bits on Mistral. Against the strongest calibrated KV-quantization baseline, HQMQ at 3.793.79 bits matches KIVI-4 (∼4.5\sim 4.5 bits) within ∼1{\sim}1 pt on CoQA, 0.60.6 pts on TruthfulQA, and 2.32.3 pts on GSM8K, at 16%16\% fewer bits and without a calibration pass. At the storage level, HQMQ delivers up to 5.05×5.05\times KV compression, shrinking a Llama-3-70B 128k-context cache from 43 GB to 8.5 GB.
Aug 31, 2026cs.CL

Faithfulness Is Not Free: Auditing Offline KV-Cache Quantization in Retrieval-Augmented Generation

Retrieval-augmented generation systems can precompute and store key-value caches of retrieved documents to avoid re-encoding context at every query. Quantizing these caches further reduces storage, but no prior work asks whether compression damages faithfulness, whether responses remain grounded in the retrieved evidence. Faithfulness and accuracy are not equivalent: a model can produce a correct answer that is no longer supported by the context it was given. We evaluate Qwen2.5-7B-Instruct under INT8 and INT4 quantization on RGB and HotpotQA, measuring both accuracy and faithfulness with a hallucination detector, NLI entailment, and an LLM judge. INT8 is near-lossless across both metrics. INT4 reduces accuracy and, more critically, even among answers that remain factually correct, over 90% of faithfulness changes are negative, i.e., accuracy metrics are blind to this regression. The harm grows under noisy retrieval and with more retrieved chunks. Faithfulness must be audited before compressed caches are deployed.
May 20, 2026cs.LG

Runtime-Certified Bounded-Error Quantized Attention

KV cache quantization reduces the memory cost of long-context LLM inference, but introduces approximation error that is typically validated only empirically. Existing systems rely on average-case robustness, with no mechanism to detect or recover from failures at runtime. We present a tiered KV cache architecture that enables runtime-certified attention: INT8 keys and INT4 values are stored in GPU memory, while FP16 originals are retained in system RAM for deterministic fallback. A two-term error decomposition yields per-head, per-step bounds on (i) attention distribution distortion from key quantization and (ii) value reconstruction error. These bounds are computed online and used to drive adaptive precision selection and a multi-stage fallback ladder, which guarantees recovery to the exact dense attention output when required. Across PG-19, NIAH, and RULER benchmarks on LLaMA~3.1-8B with contexts up to 128K, the system matches dense FP16 KV quality within noise for language modelling and retrieval tasks, while recovering catastrophic failures observed in naive INT8/INT4 baselines. Value-sensitive tasks at short context expose a controlled trade-off between compression and fidelity, which can be eliminated via tighter value tolerances or FP16-value fallback. The certification is local (per-head, per-step) and does not guarantee end-to-end model correctness, but ensures that each attention computation is either bounded relative to an FP16 reference or exactly recovered via fallback. This reframes KV cache quantization as a runtime-verified computation rather than a fixed approximation. The goal is not raw speedups, but enabling safe deployment of aggressive KV compression under strict quality constraints.