Organizations: Sharif University of Technology, Tehran, Iran · GISMA University of Applied Sciences, Potsdam, Germany · BRAINS, Brandenburg Research Center for Applied Intelligent Systems, Potsdam, Germany · Max Planck Institute for Security and Privacy, Bochum, Germany · Computer Vision Center (CVC), Barcelona, Spain · University of Cambridge, United Kingdom
Most post-training quantization pipelines fit each weight matrix to its pretrained counterpart, one matrix at a time. Whether that proxy tracks what an attention block actually computes, or how errors in the Q, K and V projections compound inside the softmax, is rarely checked. We write the objective on the attention output instead, over all three projections at once, and reuse it throughout the pipeline. JAB defines one scalar loss over the joint Q, K, V weights of a block, evaluated against the block's real causally-masked attention output, and uses it twice: to fit the quantized weights (GPTQ warm start, then STE with learnable scales), and to score the block for a multiple-choice knapsack allocation. On attention-only quantization of Mistral-7B this works. At 3 bits JAB recovers 77-90% of the gap between uniform GPTQ and full precision, and its sensitivity estimate tracks an oracle costing 73 forward passes to within a fraction of a point. It stops working once MLP layers enter the allocation. A role-aware offset rule needing no sensitivity estimate at all beats JAB on GPT-2's MLP and on the full Mistral-7B model: with a 3-bit floor it quantizes 96.4% of the weights to 4.5 bits per parameter at 6.933 perplexity, within 4.4% of full precision (6.643) at 3.56x compression, against 7.158 for JAB at the same budget. Which matrix a weight sits in matters more than any sensitivity estimate we computed. Two things came out sideways. Block-local reconstruction is an unreliable proxy for end-to-end perplexity: one run improved a block's own objective 4.6x while perplexity rose 32x, which is why every allocation here is validated end-to-end. And on attention-only quantization, fine-tuning moved weights farther from their pretrained values while pulling attention outputs closer, with net gains. Post-training seems to recover attention behavior, not weights.
Figures & tables
Pairing
fp16
Uniform
Adaptive
Gap closed
Δ Acc.
W → W
4.818
8.738
5.228
89.5%
+0.031
C4 → C4
7.492
13.506
8.898
76.6%
+0.014
W → C4
7.492
13.575
8.622
81.4%
+0.018
C4 → W
4.818
8.875
5.428
85.0%
+0.028
Table 1: Attention-only Mistral-7B at 3 bits, all four calibrate → evaluate pairings (W = WikiText-2), with the share of the uniform-versus-fp16 gap closed. Full sweep: Appendix .
B (bits)
Uniform
JAB (C1)
Oracle (C3)
3.0
53.71
120.70
46.15
4.0
26.23
27.46
26.23
Table 2: GPT-2 MLP, WikiText-2 perplexity (fp32 control 24.357 ) at the two budgets with a uniform baseline. Full sweep: Appendix , Table .
B=3.5
B=4.5
Allocator
PPL
Acc
PPL
Acc
JAB (C1)
75.009
0.296
29.041
0.394
JAB + prop (C2)
67.337
0.307
28.694
0.397
Type-offset
64.896
0.310
28.211
0.395
Table 3: GPT-2 with QKV and MLP quantized together (C4 calibration, GPTQ-only): WikiText-2 perplexity (fp32 control 24.357 ) and accuracy at the two budgets that separate the allocators. Full table: Appendix ; in-domain C4, same ranking: Appendix .
B (bits/param)
2.5
3.5
4.5
Compression
6.40×
4.57×
3.56×
Type-offset
6624.3
32.33
6.970
JAB (C1)
1799.6
8.570
7.250
JAB-norm
648.0
109.6
10.77
Table 4: Mistral-7B, all seven modules quantized: WikiText-2 perplexity (fp16 control 6.643 ) and compression of the quantized weights. Budgets B are bits per parameter (Eq. ( )). Left: criteria at fractional budgets; Type-offset needs no scoring pass. Right: uniform quantization, defined only at integer budgets, and the adaptive JAB arm. Best per budget in bold.
Refinement
Blocks refined
Blocks changed
PPL
None (GPTQ only)
–
–
7.373
Joint FT
0–31
1 (block 0)
239.52
Joint FT
0–3
1 (block 0)
239.52
Table 5: Joint refinement of the adaptive B=4.3 allocation. Block 0 is the only block whose weights change, so refining blocks 0–3 reproduces the full-depth result exactly. Full table: Appendix .
B
Compr.
JAB (C1)
JAB-norm
Type-offset
3.0
5.33×
9.713
9.713
9.713
3.5
4.57×
7.514
8.215
7.875
4.5
3.56×
7.158
7.592
6.933
Table 6: Mistral-7B, all seven modules, floor bi≥3 and per-unit traces (GPTQ-only; fp16 6.643 ; uniform 4-bit 6.938 ; adaptive B=4.37.211 ). Compression is on the quantized weights. Not comparable to Table (64 vs 128 Hessian batches); see Appendix .
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Recon. objective
Importance criterion
Same?
GPTQ ( Frantar et al., 2023 )
∥XW−XW∥2 /layer
N/A (uniform bits)
–
HAWQ-V2 ( Dong et al., 2020 )
∥XW−XW∥2 /layer
Tr(H) of task loss
No
APTQ ( Guan et al., 2024 )
Weight recon. + attn. grad.
Hessian trace, weight-space
Related
JAB-Hessian (ours)
Joint QKV attn. loss LJAB
Tr(∇2LJAB)
Identical
Fisher-coupling (ours)
Joint QKV attn. loss LJAB
∥∇LJAB∥2 at 2-bit probe
First-order
Oracle (ours)
Joint QKV attn. loss LJAB
End-to-end logit KL
Bound
Appendix
Table 7: Relationship between reconstruction objective and importance criterion. JAB closes the gap for attention; the type-offset baseline sidesteps it entirely.
Uniform
JAB (C1)
JAB + prop (C2)
Oracle (C3)
bits
GPTQ
+ FT
GPTQ
+ FT
GPTQ
+ FT
GPTQ
+ FT
2.0
5137.6
516.8
–
–
–
–
–
–
2.5
–
–
201.7
126.5
201.7
126.5
218.0
120.1
3.0
38.35
37.81
86.36
48.31
43.93
41.99
45.36
42.45
3.5
–
–
28.32
28.38 †
28.32
28.38 †
28.61
29.09
4.0
26.23
25.85
26.23
25.85
26.23
25.85
26.23
25.85
Appendix
Table 8: GPT-2 Run A: full WikiText-2 test perplexity across criteria, budgets, and modes. fp32 control 24.357 . “–” marks combinations that do not exist. Best per row in bold .
Wiki → Wiki
C4 → C4
Wiki → C4
C4 → Wiki
Method
3b
3.5b
4b
4.5b
6b
3b
3.5b
4b
4.5b
6b
3b
3.5b
4b
4.5b
6b
3b
3.5b
4b
4.5b
6b
Perplexity
Uniform (GPTQ)
8.738
–
4.855
–
4.822
13.506
–
7.571
–
7.495
13.575
–
7.580
–
7.496
8.875
–
4.870
–
4.820
Adaptive (Hutch., greedy)
5.228
5.129
4.855
4.850
4.834
8.898
8.670
7.571
7.556
7.518
8.622
8.340
7.580
7.577
7.530
5.428
5.302
4.870
4.861
4.839
Joint (uniform bits)
8.833
–
4.875
–
4.822
13.615
–
7.617
–
7.520
13.658
–
7.624
–
7.496
8.855
–
4.881
–
4.821
JAB + Joint (Hutch., greedy)
5.268
5.134
4.875
4.867
4.845
8.732
8.507
7.617
7.595
7.553
8.681
8.373
7.624
7.608
7.561
5.411
5.268
4.881
4.873
4.843
Appendix
Table 9: Mistral-7B, all four calibrate → evaluate corpus pairings. Upper block: perplexity (lower better). Lower block: next-token top-1 accuracy (higher better). fp16 baselines: PPL 4.818 (WikiText-2 eval) / 7.492 (C4 eval); Acc 0.6257 / 0.5500. “–” marks a flat-bit mode at a fractional budget, left undefined rather than silently rounded. Best per column in bold ; ties bolded.
Evaluated on WikiText-2 (cross-corpus)
Evaluated on C4 (in-domain)
Mode
bits
PPL
Acc
εA
PPL
Acc
εA
εW
uniform
3
39.338
0.3612
0.4711
54.862
0.3339
0.4621
0.3753
uniform
4
26.536
0.4038
0.2382
33.638
0.3692
0.2370
0.1739
uniform
6
24.488
0.4140
0.0695
30.955
0.3774
0.0702
0.0409
jab
3
80.140
0.2924
0.4463
99.206
0.2751
0.4287
0.4321
jab
3.5
29.425
0.3949
0.3579
36.958
0.3608
0.3504
0.2787
Appendix
Table 10: GPT-2 Run B (C4 calibration): five pipeline modes, five budgets, four metrics, evaluated on both corpora. εW is a property of the quantized weights alone and is therefore shared; the amplification factor ρ=εA/εW is plotted in Fig. (b). Controls: WikiText-2 24.357 / 0.4148; C4 30.778 / 0.3786. “–” marks a flat-bit mode at a fractional budget.
Uniform
JAB (C1)
JAB + prop (C2)
Oracle (C3)
bits
GPTQ
+ FT
GPTQ
+ FT
GPTQ
+ FT
GPTQ
+ FT
2.0
3698.3
8462.0 ‡
–
–
–
–
–
–
2.5
–
–
2458.3
1439.6
395.7
248.2
351.3
275.1
3.0
53.71
46.45
120.70
94.49
53.71
46.45
46.15
39.13
3.5
–
–
40.12
37.55
29.90
28.80
28.54
27.99
4.0
26.23
25.96
27.46
26.64
26.23
25.96
26.23
25.96
Appendix
Table 11: GPT-2 MLP: full WikiText-2 test perplexity, same pipeline, criteria and budgets as Table , with c_fc and mlp.c_proj quantized and attention left in fp32. fp32 control 24.357 ; budget is average bits per matrix over 24 units. “–” marks combinations that do not exist. Best per row in bold . Contrast the Uniform column with Table : here the criterion loses to it at both comparable budgets.
Criterion
Unit
Assignment (blocks 0–11)
PPL
Oracle (C3)
c_fc
4 4 4 4 4 4 4 4 4 4 4 4
28.54
c_proj
3 3 3 3 3 3 3 3 3 3 3 3
JAB (C1)
c_fc
4 3 3 3 3 3 4 4 4 4 4 4
40.12
c_proj
3 3 3 3 3 3 3 4 4 4 4 4
Appendix
Table 12: GPT-2 MLP, bit assignments at the 3.5-bit budget by unit and depth. The oracle’s rule is constant in depth and splits by matrix role; the Hutchinson criterion follows a depth gradient instead and pays 11.6 perplexity for it.
WikiText-2 (cross-corpus)
C4 (in-domain)
Mode
bits
PPL
Acc
εA
ρ
PPL
Acc
εA
ρ
εW
uniform
3
54.152
0.3241
0.4749
1.332
61.630
0.3035
0.4632
1.299
0.3565
uniform
4
26.423
0.4053
0.2112
1.349
34.765
0.3652
0.2068
1.321
0.1566
uniform
6
24.420
0.4149
0.0483
1.338
32.648
0.3734
0.0472
1.307
0.0361
jab
3
600.798
0.1219
0.4650
1.192
574.850
0.1188
0.4457
1.143
0.3900
jab
3.5
46.236
0.3421
0.3292
1.273
54.889
0.3177
0.3237
1.251
0.2587
Appendix
Table 13: GPT-2 MLP, C4 calibration: four pipeline modes, five budgets, evaluated on both corpora, with the amplification factor ρ=εA/εW reported per corpus ( εW is shared). Controls: WikiText-2 24.357/0.4148; C4 32.770/0.3736. Not directly comparable to Table , which is calibrated on WikiText-2.
Allocation, refinement
Blocks refined
Changed
Block-0 loss
PPL
Adaptive B=4.3 , GPTQ only
–
–
–
7.373
+ joint FT
0–31
1 (0)
1.85→0.41×10−3
239.52
+ joint FT
0–3
1 (0)
1.85→0.41×10−3
239.52
+ joint FT, λKL=0
0–31
1 (0)
MSE only
52.07
Uniform 4 , GPTQ only
–
–
–
6.998
+ joint FT
0–31
3 (0,3,4)
1.0→0.9×10−5
11.39
Appendix
Table 14: Joint fine-tuning under full coverage. “Changed” counts blocks whose quantized weights differ from the GPTQ warm start (flip rate >0 , Eq. ); “Block-0 loss” is that block’s own objective before and after refinement. Restricting refinement to blocks 0–3 reproduces the full-depth result exactly.
Allocator
B
PPL
Acc
εA
εW
ρ
Uniform
3.0
141.821
0.232
0.414
0.464
0.892
JAB
3.5
75.009
0.296
0.267
0.383
0.697
JAB+prop
3.5
67.337
0.307
0.320
0.379
0.845
Type-offset
3.5
64.896
0.310
0.421
0.303
1.390
All (flat)
4.0
29.308
0.393
0.238
0.210
1.137
JAB
4.5
29.041
0.394
0.218
0.188
1.159
Appendix
Table 15: GPT-2 with QKV and MLP quantized together (C4 calibration, GPTQ-only unless noted): WikiText-2 perplexity (fp32 control 24.357 ), accuracy, attention error εA , weight error εW , and ρ=εA/εW . C4 (in-domain, control 32.770 ) gives the same ranking; full C4 numbers are in Appendix (Table ). Best per budget in bold.
Allocator
B
PPL
Acc
εA
εW
ρ
Uniform
3.0
155.246
0.220
0.400
0.448
0.894
JAB
3.5
83.311
0.279
0.261
0.367
0.711
JAB+prop
3.5
81.408
0.280
0.316
0.371
0.852
Type-offset
3.5
78.239
0.287
0.421
0.303
1.390
All (flat)
4.0
39.053
0.352
0.237
0.205
1.160
JAB
4.5
38.626
0.355
0.219
0.183
1.195
Appendix
Table 16: GPT-2 with QKV and MLP quantized together, evaluated on C4 (in-domain; fp32 control 32.770 ), matching the WikiText-2 results of Table . Same allocators, budgets B (bits/parameter), and metrics. Best per budget in bold.
Many LLM applications require only narrow capabilities, yet standard post-training quantization (PTQ) methods allocate precision without considering the target task. This can waste bits on layers that are less relevant to the task signal while over-compressing layers that are critical for downstream behavior. We propose Task-Aware Quantization (TAQ), a training-free, weight-only mixed-precision PTQ framework that uses a small set of unlabeled task calibration prompts to allocate higher precision to task-relevant transformer layers under a fixed bit budget. TAQ estimates layer importance from hidden representations and output sensitivity, and we instantiate it with three scoring rules: TAQ-IS, based on activation information and stability; TAQ-KL, based on output-distribution sensitivity under a quantization-noise proxy; and TAQ-O, a label-informed oracle diagnostic for analyzing layer sensitivity. Across several benchmarks, TAQ outperforms task-agnostic baselines such in most settings, with especially strong gains in the accuracy--memory ratio. We further validate that these gains translate to real deployment behavior through hardware throughput and latency measurements, and analyze calibration robustness and residual-stream error propagation. Overall, TAQ turns mixed-precision PTQ from a model-centric compression step into a task-conditioned precision-allocation problem. A reference implementation is available at \href{https://anonymous.4open.science/r/TAQ-9217/README.md}{\includegraphics[height=1em]{imgs/github-mark.png}}.
Amit LeVi, Raz Lapid, Rom Himelstein +3
1Technion—Israel Institute of Technology · 2Intuit · 3Ben-Gurion University of the Negev +1
Mixed-precision quantization improves the accuracy of post-training quantization by allocating higher bitwidths to sensitive layers, but existing methods solve the allocation for a single fixed memory budget. In practice the budget varies across deployments and is unknown at calibration time. Adaptive quantization addresses this with one offline calibration that serves any budget, yet current methods score layer sensitivity in a manner that does not consider its dependency on quantization levels of other layers. We show that a layer's sensitivity depends strongly on the bitwidths of its upstream layers and that this dependence shifts the resulting preferred bit allocation. We propose MixQuant, a technique-agnostic adaptive framework that wraps any base quantizer. MixQuant marginalizes each layer's distortion over random quantized upstream configurations to obtain budget-agnostic scores, calibrates the quantizer's parameters on plans the allocator itself produces, and penalizes allocations that leave layers at the lowest bitwidths. A single greedy pass then serves any budget at deployment. Across Llama-3.2-3B, Llama-2-7B, and Mistral-7B under AWQ and GPTQ, MixQuant outperforms adaptive and mixed-precision baselines in every setting, improving average accuracy by up to 8 points and reducing perplexity from 12.43 to 10.70 at the tightest budget, while matching an ILP solver at negligible deployment cost.
Mixed-precision quantization (MPQ) has become a key technique for deploying large language models under stringent memory and compute constraints. We first identify a phenomenon that we term the Perplexity Illusion: layers ranked as important by perplexity-based sensitivity show little rank correlation with those that are most influential for complex reasoning performance, with Kendall τ≈0 in our analysis. We further reveal an Alignment-Diversity Tradeoff: using only target-task calibration data can degrade post-quantization performance, whereas incorporating general-domain data stabilizes sensitivity estimation and improves robustness across tasks. Based on these observations, we propose TASA (Task-Aware Sensitivity Analysis), a two-level framework that jointly optimizes calibration-data composition and mixed-precision bit allocation. Specifically, TASA searches for a calibration-data mixture using a training-free gradient-trace alignment criterion, and then aggregates perplexity and reasoning-oriented sensitivity signals to guide both inter-layer and intra-layer bit allocation. Experiments on LLaMA-3-8B and Qwen2.5-7B reveal a precision inversion: appropriately allocated 3.5-bit models can match or surpass less task-aware 4-bit baselines. At an average precision of 3.5 bits, TASA matches or outperforms several competitive 4-bit uniform baselines in aggregate accuracy, and improves over the strongest W3 baseline on GSM8K by more than 20 absolute points on LLaMA-3-8B. These results show that calibration-data composition substantially affects task-sensitive quantization, a factor underexplored in prior work.
Fei Wang, Chao Xue, Taoran Liu +3
South China University of Technology · Ocean University of China · Shenzhen Campus of Sun Yat-sen University +1