Organizations: Sharif University of Technology, Tehran, Iran · GISMA University of Applied Sciences, Potsdam, Germany · BRAINS, Brandenburg Research Center for Applied Intelligent Systems, Potsdam, Germany · Max Planck Institute for Security and Privacy, Bochum, Germany · Computer Vision Center (CVC), Barcelona, Spain · University of Cambridge, United Kingdom
Most post-training quantization pipelines fit each weight matrix to its pretrained counterpart, one matrix at a time. Whether that proxy tracks what an attention block actually computes, or how errors in the Q, K and V projections compound inside the softmax, is rarely checked. We write the objective on the attention output instead, over all three projections at once, and reuse it throughout the pipeline. JAB defines one scalar loss over the joint Q, K, V weights of a block, evaluated against the block's real causally-masked attention output, and uses it twice: to fit the quantized weights (GPTQ warm start, then STE with learnable scales), and to score the block for a multiple-choice knapsack allocation. On attention-only quantization of Mistral-7B this works. At 3 bits JAB recovers 77-90% of the gap between uniform GPTQ and full precision, and its sensitivity estimate tracks an oracle costing 73 forward passes to within a fraction of a point. It stops working once MLP layers enter the allocation. A role-aware offset rule needing no sensitivity estimate at all beats JAB on GPT-2's MLP and on the full Mistral-7B model: with a 3-bit floor it quantizes 96.4% of the weights to 4.5 bits per parameter at 6.933 perplexity, within 4.4% of full precision (6.643) at 3.56x compression, against 7.158 for JAB at the same budget. Which matrix a weight sits in matters more than any sensitivity estimate we computed. Two things came out sideways. Block-local reconstruction is an unreliable proxy for end-to-end perplexity: one run improved a block's own objective 4.6x while perplexity rose 32x, which is why every allocation here is validated end-to-end. And on attention-only quantization, fine-tuning moved weights farther from their pretrained values while pulling attention outputs closer, with net gains. Post-training seems to recover attention behavior, not weights.
Figures & tables
Pairing
fp16
Uniform
Adaptive
Gap closed
Δ Acc.
W → W
4.818
8.738
5.228
89.5%
+0.031
C4 → C4
7.492
13.506
8.898
76.6%
+0.014
W → C4
7.492
13.575
8.622
81.4%
+0.018
C4 → W
4.818
8.875
5.428
85.0%
+0.028
Table 1: Attention-only Mistral-7B at 3 bits, all four calibrate → evaluate pairings (W = WikiText-2), with the share of the uniform-versus-fp16 gap closed. Full sweep: Appendix .
B (bits)
Uniform
JAB (C1)
Oracle (C3)
3.0
53.71
120.70
46.15
4.0
26.23
27.46
26.23
Table 2: GPT-2 MLP, WikiText-2 perplexity (fp32 control 24.357 ) at the two budgets with a uniform baseline. Full sweep: Appendix , Table .
B=3.5
B=4.5
Allocator
PPL
Acc
PPL
Acc
JAB (C1)
75.009
0.296
29.041
0.394
JAB + prop (C2)
67.337
0.307
28.694
0.397
Type-offset
64.896
0.310
28.211
0.395
Table 3: GPT-2 with QKV and MLP quantized together (C4 calibration, GPTQ-only): WikiText-2 perplexity (fp32 control 24.357 ) and accuracy at the two budgets that separate the allocators. Full table: Appendix ; in-domain C4, same ranking: Appendix .
B (bits/param)
2.5
3.5
4.5
Compression
6.40×
4.57×
3.56×
Type-offset
6624.3
32.33
6.970
JAB (C1)
1799.6
8.570
7.250
JAB-norm
648.0
109.6
10.77
Table 4: Mistral-7B, all seven modules quantized: WikiText-2 perplexity (fp16 control 6.643 ) and compression of the quantized weights. Budgets B are bits per parameter (Eq. ( )). Left: criteria at fractional budgets; Type-offset needs no scoring pass. Right: uniform quantization, defined only at integer budgets, and the adaptive JAB arm. Best per budget in bold.
Refinement
Blocks refined
Blocks changed
PPL
None (GPTQ only)
–
–
7.373
Joint FT
0–31
1 (block 0)
239.52
Joint FT
0–3
1 (block 0)
239.52
Table 5: Joint refinement of the adaptive B=4.3 allocation. Block 0 is the only block whose weights change, so refining blocks 0–3 reproduces the full-depth result exactly. Full table: Appendix .
B
Compr.
JAB (C1)
JAB-norm
Type-offset
3.0
5.33×
9.713
9.713
9.713
3.5
4.57×
7.514
8.215
7.875
4.5
3.56×
7.158
7.592
6.933
Table 6: Mistral-7B, all seven modules, floor bi≥3 and per-unit traces (GPTQ-only; fp16 6.643 ; uniform 4-bit 6.938 ; adaptive B=4.37.211 ). Compression is on the quantized weights. Not comparable to Table (64 vs 128 Hessian batches); see Appendix .
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Recon. objective
Importance criterion
Same?
GPTQ ( Frantar et al., 2023 )
∥XW−XW∥2 /layer
N/A (uniform bits)
–
HAWQ-V2 ( Dong et al., 2020 )
∥XW−XW∥2 /layer
Tr(H) of task loss
No
APTQ ( Guan et al., 2024 )
Weight recon. + attn. grad.
Hessian trace, weight-space
Related
JAB-Hessian (ours)
Joint QKV attn. loss LJAB
Tr(∇2LJAB)
Identical
Fisher-coupling (ours)
Joint QKV attn. loss LJAB
∥∇LJAB∥2 at 2-bit probe
First-order
Oracle (ours)
Joint QKV attn. loss LJAB
End-to-end logit KL
Bound
Appendix
Table 7: Relationship between reconstruction objective and importance criterion. JAB closes the gap for attention; the type-offset baseline sidesteps it entirely.
Uniform
JAB (C1)
JAB + prop (C2)
Oracle (C3)
bits
GPTQ
+ FT
GPTQ
+ FT
GPTQ
+ FT
GPTQ
+ FT
2.0
5137.6
516.8
–
–
–
–
–
–
2.5
–
–
201.7
126.5
201.7
126.5
218.0
120.1
3.0
38.35
37.81
86.36
48.31
43.93
41.99
45.36
42.45
3.5
–
–
28.32
28.38 †
28.32
28.38 †
28.61
29.09
4.0
26.23
25.85
26.23
25.85
26.23
25.85
26.23
25.85
Appendix
Table 8: GPT-2 Run A: full WikiText-2 test perplexity across criteria, budgets, and modes. fp32 control 24.357 . “–” marks combinations that do not exist. Best per row in bold .
Wiki → Wiki
C4 → C4
Wiki → C4
C4 → Wiki
Method
3b
3.5b
4b
4.5b
6b
3b
3.5b
4b
4.5b
6b
3b
3.5b
4b
4.5b
6b
3b
3.5b
4b
4.5b
6b
Perplexity
Uniform (GPTQ)
8.738
–
4.855
–
4.822
13.506
–
7.571
–
7.495
13.575
–
7.580
–
7.496
8.875
–
4.870
–
4.820
Adaptive (Hutch., greedy)
5.228
5.129
4.855
4.850
4.834
8.898
8.670
7.571
7.556
7.518
8.622
8.340
7.580
7.577
7.530
5.428
5.302
4.870
4.861
4.839
Joint (uniform bits)
8.833
–
4.875
–
4.822
13.615
–
7.617
–
7.520
13.658
–
7.624
–
7.496
8.855
–
4.881
–
4.821
JAB + Joint (Hutch., greedy)
5.268
5.134
4.875
4.867
4.845
8.732
8.507
7.617
7.595
7.553
8.681
8.373
7.624
7.608
7.561
5.411
5.268
4.881
4.873
4.843
Appendix
Table 9: Mistral-7B, all four calibrate → evaluate corpus pairings. Upper block: perplexity (lower better). Lower block: next-token top-1 accuracy (higher better). fp16 baselines: PPL 4.818 (WikiText-2 eval) / 7.492 (C4 eval); Acc 0.6257 / 0.5500. “–” marks a flat-bit mode at a fractional budget, left undefined rather than silently rounded. Best per column in bold ; ties bolded.
Evaluated on WikiText-2 (cross-corpus)
Evaluated on C4 (in-domain)
Mode
bits
PPL
Acc
εA
PPL
Acc
εA
εW
uniform
3
39.338
0.3612
0.4711
54.862
0.3339
0.4621
0.3753
uniform
4
26.536
0.4038
0.2382
33.638
0.3692
0.2370
0.1739
uniform
6
24.488
0.4140
0.0695
30.955
0.3774
0.0702
0.0409
jab
3
80.140
0.2924
0.4463
99.206
0.2751
0.4287
0.4321
jab
3.5
29.425
0.3949
0.3579
36.958
0.3608
0.3504
0.2787
Appendix
Table 10: GPT-2 Run B (C4 calibration): five pipeline modes, five budgets, four metrics, evaluated on both corpora. εW is a property of the quantized weights alone and is therefore shared; the amplification factor ρ=εA/εW is plotted in Fig. (b). Controls: WikiText-2 24.357 / 0.4148; C4 30.778 / 0.3786. “–” marks a flat-bit mode at a fractional budget.
Uniform
JAB (C1)
JAB + prop (C2)
Oracle (C3)
bits
GPTQ
+ FT
GPTQ
+ FT
GPTQ
+ FT
GPTQ
+ FT
2.0
3698.3
8462.0 ‡
–
–
–
–
–
–
2.5
–
–
2458.3
1439.6
395.7
248.2
351.3
275.1
3.0
53.71
46.45
120.70
94.49
53.71
46.45
46.15
39.13
3.5
–
–
40.12
37.55
29.90
28.80
28.54
27.99
4.0
26.23
25.96
27.46
26.64
26.23
25.96
26.23
25.96
Appendix
Table 11: GPT-2 MLP: full WikiText-2 test perplexity, same pipeline, criteria and budgets as Table , with c_fc and mlp.c_proj quantized and attention left in fp32. fp32 control 24.357 ; budget is average bits per matrix over 24 units. “–” marks combinations that do not exist. Best per row in bold . Contrast the Uniform column with Table : here the criterion loses to it at both comparable budgets.
Criterion
Unit
Assignment (blocks 0–11)
PPL
Oracle (C3)
c_fc
4 4 4 4 4 4 4 4 4 4 4 4
28.54
c_proj
3 3 3 3 3 3 3 3 3 3 3 3
JAB (C1)
c_fc
4 3 3 3 3 3 4 4 4 4 4 4
40.12
c_proj
3 3 3 3 3 3 3 4 4 4 4 4
Appendix
Table 12: GPT-2 MLP, bit assignments at the 3.5-bit budget by unit and depth. The oracle’s rule is constant in depth and splits by matrix role; the Hutchinson criterion follows a depth gradient instead and pays 11.6 perplexity for it.
WikiText-2 (cross-corpus)
C4 (in-domain)
Mode
bits
PPL
Acc
εA
ρ
PPL
Acc
εA
ρ
εW
uniform
3
54.152
0.3241
0.4749
1.332
61.630
0.3035
0.4632
1.299
0.3565
uniform
4
26.423
0.4053
0.2112
1.349
34.765
0.3652
0.2068
1.321
0.1566
uniform
6
24.420
0.4149
0.0483
1.338
32.648
0.3734
0.0472
1.307
0.0361
jab
3
600.798
0.1219
0.4650
1.192
574.850
0.1188
0.4457
1.143
0.3900
jab
3.5
46.236
0.3421
0.3292
1.273
54.889
0.3177
0.3237
1.251
0.2587
Appendix
Table 13: GPT-2 MLP, C4 calibration: four pipeline modes, five budgets, evaluated on both corpora, with the amplification factor ρ=εA/εW reported per corpus ( εW is shared). Controls: WikiText-2 24.357/0.4148; C4 32.770/0.3736. Not directly comparable to Table , which is calibrated on WikiText-2.
Allocation, refinement
Blocks refined
Changed
Block-0 loss
PPL
Adaptive B=4.3 , GPTQ only
–
–
–
7.373
+ joint FT
0–31
1 (0)
1.85→0.41×10−3
239.52
+ joint FT
0–3
1 (0)
1.85→0.41×10−3
239.52
+ joint FT, λKL=0
0–31
1 (0)
MSE only
52.07
Uniform 4 , GPTQ only
–
–
–
6.998
+ joint FT
0–31
3 (0,3,4)
1.0→0.9×10−5
11.39
Appendix
Table 14: Joint fine-tuning under full coverage. “Changed” counts blocks whose quantized weights differ from the GPTQ warm start (flip rate >0 , Eq. ); “Block-0 loss” is that block’s own objective before and after refinement. Restricting refinement to blocks 0–3 reproduces the full-depth result exactly.
Allocator
B
PPL
Acc
εA
εW
ρ
Uniform
3.0
141.821
0.232
0.414
0.464
0.892
JAB
3.5
75.009
0.296
0.267
0.383
0.697
JAB+prop
3.5
67.337
0.307
0.320
0.379
0.845
Type-offset
3.5
64.896
0.310
0.421
0.303
1.390
All (flat)
4.0
29.308
0.393
0.238
0.210
1.137
JAB
4.5
29.041
0.394
0.218
0.188
1.159
Appendix
Table 15: GPT-2 with QKV and MLP quantized together (C4 calibration, GPTQ-only unless noted): WikiText-2 perplexity (fp32 control 24.357 ), accuracy, attention error εA , weight error εW , and ρ=εA/εW . C4 (in-domain, control 32.770 ) gives the same ranking; full C4 numbers are in Appendix (Table ). Best per budget in bold.
Allocator
B
PPL
Acc
εA
εW
ρ
Uniform
3.0
155.246
0.220
0.400
0.448
0.894
JAB
3.5
83.311
0.279
0.261
0.367
0.711
JAB+prop
3.5
81.408
0.280
0.316
0.371
0.852
Type-offset
3.5
78.239
0.287
0.421
0.303
1.390
All (flat)
4.0
39.053
0.352
0.237
0.205
1.160
JAB
4.5
38.626
0.355
0.219
0.183
1.195
Appendix
Table 16: GPT-2 with QKV and MLP quantized together, evaluated on C4 (in-domain; fp32 control 32.770 ), matching the WikiText-2 results of Table . Same allocators, budgets B (bits/parameter), and metrics. Best per budget in bold.