Quantization-aware training (QAT) leverages lower-precision arithmetic to reduce the cost of LLM deployment, but aggressive quantization degrades final model performance. A common remedy is mixed-precision training, in which high precision is assigned to some of the layers to maintain performance while keeping the cost constrained. This approach then requires precision assignments for model layers during training. We provide a new approach, called Q-PACE, consisting of a second-order sensitivity model that predicts the loss increase as a sum of quantization noise MSE weighted by per-layer curvature coefficients. During training, we periodically re-compute these coefficients using perturbations across layers, and re-assign precision. Pretraining and supervised fine-tuning experiments on LLMs of up to 4B parameters show that Q-PACE consistently improves over existing mixed-precision training recipes, and achieves comparable loss at substantially lower total memory budgets. We further find that quantization sensitivity is highly predictable by depth and layer type, and its stability during training allows for infrequent, cheap recalibration.
Figures & tables
Figure 1: Final loss improvement for a 500M model trained with mixed precision across methods and budgets. We improve over last layers in BF16 ( NVIDIA et al., 2026 ) and SNIP ( Pan et al., 2026 ) .
Figure 2: Test loss vs. the modeled sensitivity-weighted forward error for 500M models’ final allocations.
Model
Budget
FP4
FP8
NVIDIA
Rec.
SNIP
Rec.
Q-PACE
Rec.
gain
350 M, 16 B tok
5.0{4,6,8,16}
2.886
2.848
2.884
6%
2.872
37%
2.868
49%
7.5×
5.5{4,6,8,16}
2.880
15%
2.867
51%
2.858
74%
4.8×
6.5{4,8,16}
2.876
28%
2.864
59%
2.857
75%
2.7×
6.5{4,16}
2.876
28%
2.874
33%
2.870
43%
1.6×
500 M, 74 B tok
5.0{4,6,8,16}
2.693
2.657
2.692
2%
2.678
40%
2.673
55%
32.4×
5.5{4,6,8,16}
2.688
14%
2.673
54%
2.664
80%
5.8×
Table 1: Final test loss across models. Budget equals target average bits per parameter. “Rec.” (recovery) is the share of the FP4-to-FP8 gap recovered, (LFP4−L)/(LFP4−LFP8) . The column “gain” reports the ratio between the gain of Q-PACE over FP4 and the gain of the NVIDIA allocation over FP4, (LFP4−LQ-PACE )/(LFP4−LNVIDIA) . The std error measured over 3 random seeds for 350M model is ≤0.001 . Best loss per row is shown in bold.
Figure 4
Figure 4: Pretraining experiment for a 4B model. Test-loss improvement over uniform FP4; labels give the recovered gap between FP4 and FP8 trainings, realized bits/parameter, and the precision options in parentheses.
Figure 6Figure 7
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
AdamW
Muon
Model
Budget
ε=MSE
ε′=∥Wi∥F2∥δWi∥F2
Δ
ε=MSE
ε′=∥Wi∥F2∥δWi∥F2
Δ
350M
6.0
2.8598 ± .0014
2.8600 ± .0003
+0.0002
2.8434 ± .0018
2.8392 ± .0014
−0.0041
800M
5.0
2.6122
2.6124
+0.0002
2.6105
2.6090
−0.0015
6.0
2.6071
2.6087
+0.0016
2.6080
2.6022
−0.0058
Appendix
Table 3: Final test loss of Q-PACE mixed-precision runs with the error measured as absolute MSE vs. relative to the weight norm; Δ=L(ε′)−L(ε) .
Model
Budget
Error
Final test loss
Improvement over FP4
350M AdamW
6.33
Analytic
2.8598±0.0014
0.527±0.049%
Measured
2.8598±0.0017
0.527±0.060%
800M AdamW
5.39
Analytic
2.6081
0.298%
Measured
2.6095
0.242%
Appendix
Table 4: Q-PACE with the analytic vs the measured quantization error for MSE estimation in Eq. ( 5 ) model.
Figure 7: Loss increase from quantization, L−Lbf16 (test loss minus that of the same weights in BF16), against the sensitivity-weighted weight error ∑lαlel of each model’s final allocation, with αl refitted on its final weights. Left: fwd + bwd mixed; right: fwd mixed, bwd FP4.
Table 5: Pretraining loss and recovered share of the NVFP4-to-MXFP8 gap (rec.) per budget cell when the tiers set both passes and when they set the forward only (FP4 backward everywhere; each setting against its own MXFP8 run).
350M
500M
800M
4B
blocks / heads / GQA ratio (KV heads)
11 / 12 / 4 (3)
17 / 12 / 3 (4)
17 / 16 / 4 (4)
28 / 28 / 4 (7)
hidden / head dim
1536 / 128
1536 / 128
2048 / 128
3584 / 128
parameters (total / non-embedding)
370.8M / 272.5M
526.2M / 427.9M
814.3M / 748.8M
4.006B / 3.777B
switchable linears
66
102
102
168
context
1024
1024
1024
4096
global batch (seqs / tokens)
512 / 524k
512 / 524k
1024 / 1.05M
256 / 1.05M
Appendix
Table 6: Pretraining configurations.
Figure 8: Test loss gain over uniform FP4 for 350M model over the number of noise levels m per layer and calibration sequences ∣D∣ . Performance improves over the smallest calibration setting (1,1) and then remains relatively stable across a broad range of m and ∣D∣ , indicating diminishing returns from additional calibration compute.
Method
Final validation loss
Improvement over FP4
Q-PACE inversed
2.8551±0.0010
0.554±0.038%
Q-PACE
2.8554±0.0011
0.544±0.037%
Q-PACE: measured MSE( W )
2.8569±0.0021
0.491±0.061%
Q-PACE: MSE( Wgrad )
2.8570±0.0005
0.488±0.031%
Quadratic Q-PACE
2.8575±0.0018
0.472±0.050%
Q-PACE: weights + activations MSE
2.8581±0.0010
0.450±0.048%
Appendix
Table 7: Test loss results for 350M model trained on 18B tokens with budget 6.5 bit/parameter and choices 4,16. Average and std error over 3 random seeds.
Figure 9: Share of different layer types promoted to BF16 over the training duration.
Figure 10: Relative sensitivity by layer type over training for the 800M model. Values are averaged across layers of the same type at each calibration point. Q-PACE assigns the largest relative sensitivity to the MLP projections, whereas SNIP places comparatively greater sensitivity on K projections.
Figure 11: Relative sensitivity over training for the 800M model.
Figure 12: Bitwidth allocations throughout training for 800M model at budget 5.5 bits/parameter with bit choices 4,16.
Figure 13: Layer-wise sensitivity heatmap for Olmo-3-7B ( Olmo: et al., 2026 ) during training. Lighter color corresponds to higher sensitivity α .
Figure 14: Loss improvement over uniform NVFP4 against the realized bits per weight of the trained allocation, for pretraining at 350M and 500M and for SFT.
Figure 15: Normalized sensitivity of Q-PACE and SNIP at budget 6.5 on {4,8,16} at 350M and 500M. Type shares over the calibrations, depth shares, and the final type-by-block map.
Figure 16: SFT training loss and gradient norm (mean of three seeds) and the normalized sensitivity by projection type of Q-PACE and SNIP at budget 6.5 on {4,8,16} .
Figure 17: 500M allocations of Q-PACE and SNIP at budget 6.5 on {4,8,16} over training. Parameter share per format and bits per weight by projection type.
Figure 18: Final 350M allocations of Q-PACE and SNIP for every budget and menu.
Figure 19: Final 500M allocations of Q-PACE and SNIP for every budget and menu.
Figure 20: Final SFT allocations of Q-PACE and SNIP for every budget and menu. Qwen3 has three MLP projections per block.
Figure 21: 4B normalized sensitivity at budget 6.5 on {4,16} . Type shares over the calibrations, depth shares, and the latest type-by-block map.
Figure 22: 4B sensitivity of all 168 layers across the calibrations at budget 6.5 on {4,16} . Colors are clipped at the 99th percentile, on separate scales for the two estimators.
Figure 23: 4B final allocations and their dynamics. Parameter share per format and bits per weight by projection type.
Figure 24: Final sensitivity (logarithmic color scale) and format of all 168 switchable layers in the four 4B Q-PACE runs.
run
GSM8K
IFEval
HumanEval
NQ
TQA
uniform references
uniform NVFP4
80.6±1.4
30.6±0.8
69.1±2.9
8.7±0.4
51.2±0.7
uniform MXFP8
84.2±0.3
34.3±1.1
75.8±0.9
11.2±0.7
52.6±0.1
budget 5, menu {4,6,8,16}
NVIDIA
81.0±0.9
28.0±1.7
71.3±2.7
9.2±0.5
51.2±0.2
Q-PACE
81.1±0.7
33.0±2.2
72.8±3.1
8.9±0.6
49.9±0.6
Appendix
Table 8: SFT generative and instruction benchmarks (accuracy %, mean ± seed standard deviation over three seeds).
run
ARC-c
ARC-e
HellaSw
PIQA
Wino
LAMBADA
MMLU
uniform references
uniform NVFP4
44.9±0.6
74.6±0.3
53.8±0.1
76.1±0.4
67.6±0.8
66.3±0.2
68.7±0.1
uniform MXFP8
47.0±0.2
78.4±0.1
54.6±0.0
78.0±0.1
70.7±0.6
68.6±0.1
71.1±0.1
budget 5, menu {4,6,8,16}
NVIDIA
44.6±0.5
74.6±0.7
53.7±0.2
76.1±0.3
67.9±1.2
67.0±0.2
68.8±0.1
Q-PACE
46.3±0.7
76.9±0.1
53.9±0.1
76.3±0.3
68.4±0.2
67.2±0.2
69.5±0.1
Appendix
Table 9: SFT knowledge/likelihood tasks (accuracy %, mean ± seed standard deviation over three seeds).
run
BoolQ
Copa
RACE
LogiQA
ASDiv
SAT-Math
NQ
uniform references
uniform NVFP4
77.1±0.2
84.0±1.7
39.5±0.7
40.2±0.7
27.9±0.2
43.6±2.0
8.7±0.4
uniform MXFP8
81.2±0.3
83.0±0.0
40.5±0.1
42.8±0.6
26.6±0.5
48.5±2.1
11.2±0.7
budget 5, menu {4,6,8,16}
NVIDIA
79.3±0.4
83.7±0.6
40.4±0.9
39.6±0.8
27.4±0.9
42.7±1.4
9.2±0.5
Q-PACE
81.9±0.6
83.7±3.1
40.1±0.9
41.0±0.8
28.5±1.0
45.8±1.3
8.9±0.6
Appendix
Table 10: SFT probe tasks (accuracy %, mean ± seed standard deviation over three seeds).
Many LLM applications require only narrow capabilities, yet standard post-training quantization (PTQ) methods allocate precision without considering the target task. This can waste bits on layers that are less relevant to the task signal while over-compressing layers that are critical for downstream behavior. We propose Task-Aware Quantization (TAQ), a training-free, weight-only mixed-precision PTQ framework that uses a small set of unlabeled task calibration prompts to allocate higher precision to task-relevant transformer layers under a fixed bit budget. TAQ estimates layer importance from hidden representations and output sensitivity, and we instantiate it with three scoring rules: TAQ-IS, based on activation information and stability; TAQ-KL, based on output-distribution sensitivity under a quantization-noise proxy; and TAQ-O, a label-informed oracle diagnostic for analyzing layer sensitivity. Across several benchmarks, TAQ outperforms task-agnostic baselines such in most settings, with especially strong gains in the accuracy--memory ratio. We further validate that these gains translate to real deployment behavior through hardware throughput and latency measurements, and analyze calibration robustness and residual-stream error propagation. Overall, TAQ turns mixed-precision PTQ from a model-centric compression step into a task-conditioned precision-allocation problem. A reference implementation is available at \href{https://anonymous.4open.science/r/TAQ-9217/README.md}{\includegraphics[height=1em]{imgs/github-mark.png}}.
Amit LeVi, Raz Lapid, Rom Himelstein +3
1Technion—Israel Institute of Technology · 2Intuit · 3Ben-Gurion University of the Negev +1
Mixed-precision quantization improves the accuracy of post-training quantization by allocating higher bitwidths to sensitive layers, but existing methods solve the allocation for a single fixed memory budget. In practice the budget varies across deployments and is unknown at calibration time. Adaptive quantization addresses this with one offline calibration that serves any budget, yet current methods score layer sensitivity in a manner that does not consider its dependency on quantization levels of other layers. We show that a layer's sensitivity depends strongly on the bitwidths of its upstream layers and that this dependence shifts the resulting preferred bit allocation. We propose MixQuant, a technique-agnostic adaptive framework that wraps any base quantizer. MixQuant marginalizes each layer's distortion over random quantized upstream configurations to obtain budget-agnostic scores, calibrates the quantizer's parameters on plans the allocator itself produces, and penalizes allocations that leave layers at the lowest bitwidths. A single greedy pass then serves any budget at deployment. Across Llama-3.2-3B, Llama-2-7B, and Mistral-7B under AWQ and GPTQ, MixQuant outperforms adaptive and mixed-precision baselines in every setting, improving average accuracy by up to 8 points and reducing perplexity from 12.43 to 10.70 at the tightest budget, while matching an ILP solver at negligible deployment cost.
Mixed-precision quantization (MPQ) has become a key technique for deploying large language models under stringent memory and compute constraints. We first identify a phenomenon that we term the Perplexity Illusion: layers ranked as important by perplexity-based sensitivity show little rank correlation with those that are most influential for complex reasoning performance, with Kendall τ≈0 in our analysis. We further reveal an Alignment-Diversity Tradeoff: using only target-task calibration data can degrade post-quantization performance, whereas incorporating general-domain data stabilizes sensitivity estimation and improves robustness across tasks. Based on these observations, we propose TASA (Task-Aware Sensitivity Analysis), a two-level framework that jointly optimizes calibration-data composition and mixed-precision bit allocation. Specifically, TASA searches for a calibration-data mixture using a training-free gradient-trace alignment criterion, and then aggregates perplexity and reasoning-oriented sensitivity signals to guide both inter-layer and intra-layer bit allocation. Experiments on LLaMA-3-8B and Qwen2.5-7B reveal a precision inversion: appropriately allocated 3.5-bit models can match or surpass less task-aware 4-bit baselines. At an average precision of 3.5 bits, TASA matches or outperforms several competitive 4-bit uniform baselines in aggregate accuracy, and improves over the strongest W3 baseline on GSM8K by more than 20 absolute points on LLaMA-3-8B. These results show that calibration-data composition substantially affects task-sensitive quantization, a factor underexplored in prior work.
Fei Wang, Chao Xue, Taoran Liu +3
South China University of Technology · Ocean University of China · Shenzhen Campus of Sun Yat-sen University +1