Shared Stopping Decisions Change Answers in HQQ Cache Quantization
Organizations: Department of Artificial Intelligence, Jeju National University · School of Software, Soongsil University
Abstract
Language-model systems batch questions for throughput, but unrelated questions should not change a target's answer when its input and numerical execution are fixed. We study compression of the key and value cache, which stores attention representations reused during generation. With request-local groups, Transformers' Half-Quadratic Quantization (HQQ) backend updates compression parameters separately but uses a shared average error to decide when all updates stop. Replacing only the question batched with the target changes four-bit HQQ answers in 170/384 test comparisons across two models. Replaying the other execution's update counts reproduces its complete answer and cache fingerprints in every changed pair, in both directions. Computing the stopping mean in FP32 reduces cache differences but leaves answer changes. Native HQQ also changes confirmed numerical correctness in eight arithmetic pairs. Fixed iterations and request-local stopping remove observed companion dependence under matched controls. Request-local stopping remains sensitive to synthetic padding changes at the tensor level. Fixing the original iteration budget removes this decision path without tuning. Neither repair has an established quality advantage, and natural rebatching still changes answers. Request-independence audits must cover stopping decisions as well as quantization groups.
Figures & tables
| Model | Task | Text | Numerical correctness /resolved |
|---|---|---|---|
| Llama | Math | 80/128 | 6/122 |
| Qwen | Math | 89/128 | 2/124 |
| Both | Math | 169/256 | 8/246 |
| Llama | QA | 0/64 | – |
| Qwen | QA | 1/64 | – |
| Method | Text changes | Correctness changes |
|---|---|---|
| BF16 | 36.1% | 17/1,484 |
| Native HQQ | 54.9% | 64/1,501 |
| Fixed 20 | 46.6% | 62/1,496 |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Method | GSM | IFEval | Hotpot | MuSiQue | SQuAD | TriviaQA |
|---|---|---|---|---|---|---|---|
| Llama | BF16 | 73.96 | 83.33 | 87.92 | 34.32 | 85.72 | 91.53 |
| Llama | Native HQQ | 66.67 | 81.25 | 87.92 | 34.32 | 85.72 | 91.53 |
| Llama | No optimization | 67.71 | 82.29 | 87.92 | 34.32 | 85.72 | 91.53 |
| Llama | Fixed 2 | 68.75 | 83.33 | 87.92 | 33.98 | 85.72 | 91.53 |
| Llama | Fixed 4 | 65.62 | 81.25 | 87.92 | 33.98 | 85.72 | 91.53 |
| Llama | Fixed 8 | 67.71 | 83.33 | 87.92 | 33.98 | 85.72 | 91.53 |
| Model | 2b text | 2b pass | 4b text | 4b pass |
|---|---|---|---|---|
| Llama 3.1 8B | 71 | 12 | 49 | 9 |
| Qwen3 8B | 126 | 8 | 40 | 3 |
| Mistral 7B | 80 | 2 | 53 | 4 |
| Phi-4 | 83 | 7 | 46 | 0 |
| OLMo 2 1B | 44 | 5 | 39 | 0 |
| Dataset | 2b text | 2b pass | 4b text | 4b pass |
|---|---|---|---|---|
| GSM8K | 142 | 6 | 92 | 12 |
| HotpotQA | 32 | 2 | 2 | 0 |
| IFEval | 146 | 22 | 121 | 4 |
| MuSiQue | 44 | 0 | 7 | 0 |
| SQuAD | 20 | 2 | 2 | 0 |
| TriviaQA | 20 | 2 | 3 | 0 |
| Initial pairs | Initial calls | Different counts | Different caches | Changed answers / generated pairs | ||||
|---|---|---|---|---|---|---|---|---|
| Model/task | FP16 | FP32 | FP16 | FP32 | FP16 | FP32 | ||
| Llama/math | 128 | 8,192 | 3,900 | 1,872 | 3,900 | 272 | 80/128 | 62/128 |
| Llama/QA | 64 | 4,096 | 1,473 | 316 | 1,473 | 160 | 0/16 | 0/16 |
| Qwen/math | 128 | 9,216 | 3,430 | 806 | 3,430 | 101 | 89/128 | 67/128 |
| Qwen/QA | 64 | 4,608 | 1,740 | 147 | 1,740 | 81 | 0/16 | 0/16 |
| median (%) | median (%) | |||
|---|---|---|---|---|
| Model/task | FP16 | FP32 | FP16 | FP32 |
| Llama/math | 0.687 | 0.108 | 10.34 | 10.34 |
| Llama/QA | 0.576 | 0.035 | 10.48 | 10.48 |
| Qwen/math | 1.084 | 0.167 | 11.45 | 11.43 |
| Qwen/QA | 1.018 | 0.036 | 11.02 | 11.02 |
| Different initial caches | Llama | Qwen |
|---|---|---|
| 0 | 1/13 | 22/54 |
| 1–4 | 57/107 | 45/74 |
| 5–16 | 4/8 | – |
| Model/task | Error | |||||
|---|---|---|---|---|---|---|
| Llama/math | (%) | 10.7628 | 10.5295 | 10.4160 | 10.3455 | 10.3371 |
| Llama/math | MAE ( ) | 55.1617 | 54.0878 | 53.5362 | 53.0082 | 52.7284 |
| Llama/QA | (%) | 10.9478 | 10.7150 | 10.5895 | 10.5285 | 10.5193 |
| Llama/QA | MAE ( ) | 57.6475 | 56.5702 | 55.9558 | 55.4022 | 55.1108 |
| Qwen/math | (%) | 12.0867 | 11.8574 | 11.7241 | 11.6003 | 11.4400 |
| Qwen/math | MAE ( ) | 188.9711 | 184.7405 | 181.8135 | 178.8776 | 176.2468 |
| Model | GSM8K test ID | Assertion A | Assertion B | Reference | A to B |
|---|---|---|---|---|---|
| Llama | 439 | #### 17 | #### 7 | 17 | correct wrong |
| Llama | 1310 | #### 64 | #### 32 | 64 | correct wrong |
| Llama | 1092 | #### 2160 | #### 1080 | 1080 | wrong correct |
| Llama | 406 | #### 200 | #### 120 | 200 | correct wrong |
| Llama | 532 | #### 25 | #### 35 | 25 | correct wrong |
| Llama | 648 | #### 56 | #### 54 | 54 | wrong correct |
| Pass changes | Same | Different | Unresolved |
|---|---|---|---|
| Yes | 5 | 6 | 5 |
| No | 232 | 3 | 5 |
| Llama | Qwen | |||||
|---|---|---|---|---|---|---|
| Method | GSM strict | Final value | Hotpot F1 | GSM strict | Final value | Hotpot F1 |
| BF16 | 78.12 | 89.06–90.62 | 70.71 | 66.41 | 90.62–96.09 | 65.08 |
| Native | 74.61 | 87.11–90.23 | 69.82 | 69.14 | 93.36–95.70 | 62.69 |
| Unoptimized | 71.88 | 83.59–87.50 | 69.82 | 64.84 | 91.41–95.31 | 65.02 |
| Fixed 4 | 75.00 | 86.72–90.62 | 69.82 | 66.41 | 92.97–96.09 | 65.56 |
| Fixed 20 | 75.00 | 87.50–91.41 | 69.82 | 67.19 | 93.75–96.09 | 62.69 |
| Dataset | Model | Text | Pass | Correct /all | |
|---|---|---|---|---|---|
| GSM8K | Llama | 128 | 62.50 [53.91, 71.09] | 7.03 [3.12, 11.72] | 4.69 [1.56, 8.59] |
| GSM8K | Qwen | 128 | 69.53 [61.72, 77.34] | 5.47 [2.34, 10.16] | 1.56 [0.00, 3.91] |
| HotpotQA | Llama | 64 | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | – |
| HotpotQA | Qwen | 64 | 1.56 [0.00, 4.69] | 0.00 [0.00, 0.00] | – |
| Native update count | Complete compression (ms) | |||||
|---|---|---|---|---|---|---|
| Model/task | Median | Range | Below 20 | Native | Fixed 20 | No optimization |
| Llama/math | 13 | 9–20 | 20/24 | 4.46 | 4.11 | 0.48 |
| Llama/QA | 14 | 10–20 | 16/24 | 7.58 | 8.75 | 0.38 |
| Qwen/math | 19.5 | 9–20 | 12/24 | 3.93 | 3.56 | 0.34 |
| Qwen/QA | 20 | 10–20 | 11/24 | 12.25 | 11.69 | 1.18 |
| All cells | 18 | 9–20 | 59/96 | 6.35 | 6.22 | 0.52 |
| Method | Time (ms) |
|---|---|
| Native | 8.3345 |
| Unoptimized | 1.2849 |
| Fixed 2 | 2.5221 |
| Fixed 4 | 3.1205 |
| Fixed 8 | 4.3243 |
| Fixed 20 | 7.8219 |
| Shuffled | Sorted | |||
|---|---|---|---|---|
| Method | Long | Short | Long | Short |
| BF16 | 74 | 1 | 82 | 3 |
| Native HQQ | 115 | 5 | 111 | 11 |
| No optimization | 97 | 6 | 93 | 8 |
| Fixed 2 | 92 | 5 | 93 | 10 |
| Fixed 4 | 94 | 7 | 98 | 11 |
| Shuffled | Sorted | |||
|---|---|---|---|---|
| Method | Strict | Num. | Strict | Num. |
| BF16 | 2 | 2/61 | 6 | 2/62 |
| Native HQQ | 8 | 3/61 | 3 | 3/62 |
| No optimization | 7 | 1/60 | 8 | 1/59 |
| Fixed 2 | 5 | 1/60 | 8 | 4/60 |
| Fixed 4 | 7 | 1/60 | 10 | 2/63 |
| Changed text / 384 pairs | Math correctness changes / resolved pairs | |||||
| Shuffle seed | BF16 | Native | Fixed 20 | BF16 | Native | Fixed 20 |
| 1 | 136 | 218 | 168 | 3/248 | 12/248 | 9/249 |
| 2 | 130 | 218 | 184 | 3/248 | 12/250 | 9/247 |
| 3 | 144 | 215 | 190 | 3/247 | 10/251 | 11/249 |
| 4 | 141 | 205 | 183 | 3/247 | 11/252 | 9/250 |
| 5 | 138 | 197 | 175 | 3/248 | 10/249 | 11/251 |
| Model/task | Measure | BF16 | Native | Fixed 20 |
|---|---|---|---|---|
| Llama/math | Text | 50.91 | 79.69 | 69.27 |
| Numeric bounds | 88.18–88.87 | 87.79–89.65 | 87.21–89.45 | |
| Llama/QA | Text | 1.30 | 5.47 | 4.69 |
| EM/F1 | 55.47/70.24 | 52.54/68.94 | 52.34/69.08 | |
| Qwen/math | Text | 54.95 | 79.69 | 65.62 |
| Numeric bounds | 90.92–96.48 | 93.65–96.09 | 94.24–96.09 |
| Numerical outcome | Llama | Qwen | Mistral |
|---|---|---|---|
| Same correct | 202 | 224 | 29 |
| Correct wrong | 4 | 2 | 2 |
| Wrong correct | 3 | 1 | 5 |
| Different wrong | 5 | 1 | 12 |
| Same wrong | 22 | 13 | 8 |
| Unresolved | 20 | 15 | 200 |
| Model | Method | Final math numbers | Math official | QA EM | QA F1 | Length limit |
|---|---|---|---|---|---|---|
| Llama | BF16 | 221/28/7 | 174 | 55.47 | 71.02 | 5 |
| Native separate | 208/36/12 | 160 | 56.25 | 71.06 | 11 | |
| Native joint | 208/35/13 | 161 | 56.25 | 71.06 | 10 | |
| No optimization | 209/36/11 | 165 | 56.25 | 70.38 | 10 | |
| Fixed 2 | 209/35/12 | 161 | 57.03 | 70.10 | 11 | |
| Fixed 4 | 219/27/10 | 169 | 55.47 | 69.47 | 8 |
| Study | Model | Task | Median length | Full changes | First 64 | After 64 only |
|---|---|---|---|---|---|---|
| Cache | Llama | Math | 202 | 80/128 | 39/128 | 41/128 |
| Cache | Llama | QA | 5 | 0/64 | 0/64 | 0/64 |
| Cache | Qwen | Math | 266.5 | 89/128 | 36/128 | 53/128 |
| Cache | Qwen | QA | 5 | 1/64 | 1/64 | 0/64 |
| Weight | Llama | Math | 216 | 162/256 | 80/256 | 82/256 |
| Weight | Llama | QA | 5 | 7/128 | 7/128 | 0/128 |