Quantization-aware training (QAT) leverages lower-precision arithmetic to reduce the cost of LLM deployment, but aggressive quantization degrades final model performance. A common remedy is mixed-precision training, in which high precision is assigned to some of the layers to maintain performance while keeping the cost constrained. This approach then requires precision assignments for model layers during training. We provide a new approach, called Q-PACE, consisting of a second-order sensitivity model that predicts the loss increase as a sum of quantization noise MSE weighted by per-layer curvature coefficients. During training, we periodically re-compute these coefficients using perturbations across layers, and re-assign precision. Pretraining and supervised fine-tuning experiments on LLMs of up to 4B parameters show that Q-PACE consistently improves over existing mixed-precision training recipes, and achieves comparable loss at substantially lower total memory budgets. We further find that quantization sensitivity is highly predictable by depth and layer type, and its stability during training allows for infrequent, cheap recalibration.
Figures & tables
Figure 1: Final loss improvement for a 500M model trained with mixed precision across methods and budgets. We improve over last layers in BF16 ( NVIDIA et al., 2026 ) and SNIP ( Pan et al., 2026 ) .
Figure 2: Test loss vs. the modeled sensitivity-weighted forward error for 500M models’ final allocations.
Model
Budget
FP4
FP8
NVIDIA
Rec.
SNIP
Rec.
Q-PACE
Rec.
gain
350 M, 16 B tok
5.0{4,6,8,16}
2.886
2.848
2.884
6%
2.872
37%
2.868
49%
7.5×
5.5{4,6,8,16}
2.880
15%
2.867
51%
2.858
74%
4.8×
6.5{4,8,16}
2.876
28%
2.864
59%
2.857
75%
2.7×
6.5{4,16}
2.876
28%
2.874
33%
2.870
43%
1.6×
500 M, 74 B tok
5.0{4,6,8,16}
2.693
2.657
2.692
2%
2.678
40%
2.673
55%
32.4×
5.5{4,6,8,16}
2.688
14%
2.673
54%
2.664
80%
5.8×
Table 1: Final test loss across models. Budget equals target average bits per parameter. “Rec.” (recovery) is the share of the FP4-to-FP8 gap recovered, (LFP4−L)/(LFP4−LFP8) . The column “gain” reports the ratio between the gain of Q-PACE over FP4 and the gain of the NVIDIA allocation over FP4, (LFP4−LQ-PACE )/(LFP4−LNVIDIA) . The std error measured over 3 random seeds for 350M model is ≤0.001 . Best loss per row is shown in bold.
Figure 4
Figure 4: Pretraining experiment for a 4B model. Test-loss improvement over uniform FP4; labels give the recovered gap between FP4 and FP8 trainings, realized bits/parameter, and the precision options in parentheses.
Figure 6Figure 7
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
AdamW
Muon
Model
Budget
ε=MSE
ε′=∥Wi∥F2∥δWi∥F2
Δ
ε=MSE
ε′=∥Wi∥F2∥δWi∥F2
Δ
350M
6.0
2.8598 ± .0014
2.8600 ± .0003
+0.0002
2.8434 ± .0018
2.8392 ± .0014
−0.0041
800M
5.0
2.6122
2.6124
+0.0002
2.6105
2.6090
−0.0015
6.0
2.6071
2.6087
+0.0016
2.6080
2.6022
−0.0058
Appendix
Table 3: Final test loss of Q-PACE mixed-precision runs with the error measured as absolute MSE vs. relative to the weight norm; Δ=L(ε′)−L(ε) .
Model
Budget
Error
Final test loss
Improvement over FP4
350M AdamW
6.33
Analytic
2.8598±0.0014
0.527±0.049%
Measured
2.8598±0.0017
0.527±0.060%
800M AdamW
5.39
Analytic
2.6081
0.298%
Measured
2.6095
0.242%
Appendix
Table 4: Q-PACE with the analytic vs the measured quantization error for MSE estimation in Eq. ( 5 ) model.
Figure 7: Loss increase from quantization, L−Lbf16 (test loss minus that of the same weights in BF16), against the sensitivity-weighted weight error ∑lαlel of each model’s final allocation, with αl refitted on its final weights. Left: fwd + bwd mixed; right: fwd mixed, bwd FP4.
Table 5: Pretraining loss and recovered share of the NVFP4-to-MXFP8 gap (rec.) per budget cell when the tiers set both passes and when they set the forward only (FP4 backward everywhere; each setting against its own MXFP8 run).
350M
500M
800M
4B
blocks / heads / GQA ratio (KV heads)
11 / 12 / 4 (3)
17 / 12 / 3 (4)
17 / 16 / 4 (4)
28 / 28 / 4 (7)
hidden / head dim
1536 / 128
1536 / 128
2048 / 128
3584 / 128
parameters (total / non-embedding)
370.8M / 272.5M
526.2M / 427.9M
814.3M / 748.8M
4.006B / 3.777B
switchable linears
66
102
102
168
context
1024
1024
1024
4096
global batch (seqs / tokens)
512 / 524k
512 / 524k
1024 / 1.05M
256 / 1.05M
Appendix
Table 6: Pretraining configurations.
Figure 8: Test loss gain over uniform FP4 for 350M model over the number of noise levels m per layer and calibration sequences ∣D∣ . Performance improves over the smallest calibration setting (1,1) and then remains relatively stable across a broad range of m and ∣D∣ , indicating diminishing returns from additional calibration compute.
Method
Final validation loss
Improvement over FP4
Q-PACE inversed
2.8551±0.0010
0.554±0.038%
Q-PACE
2.8554±0.0011
0.544±0.037%
Q-PACE: measured MSE( W )
2.8569±0.0021
0.491±0.061%
Q-PACE: MSE( Wgrad )
2.8570±0.0005
0.488±0.031%
Quadratic Q-PACE
2.8575±0.0018
0.472±0.050%
Q-PACE: weights + activations MSE
2.8581±0.0010
0.450±0.048%
Appendix
Table 7: Test loss results for 350M model trained on 18B tokens with budget 6.5 bit/parameter and choices 4,16. Average and std error over 3 random seeds.
Figure 9: Share of different layer types promoted to BF16 over the training duration.
Figure 10: Relative sensitivity by layer type over training for the 800M model. Values are averaged across layers of the same type at each calibration point. Q-PACE assigns the largest relative sensitivity to the MLP projections, whereas SNIP places comparatively greater sensitivity on K projections.
Figure 11: Relative sensitivity over training for the 800M model.
Figure 12: Bitwidth allocations throughout training for 800M model at budget 5.5 bits/parameter with bit choices 4,16.
Figure 13: Layer-wise sensitivity heatmap for Olmo-3-7B ( Olmo: et al., 2026 ) during training. Lighter color corresponds to higher sensitivity α .
Figure 14: Loss improvement over uniform NVFP4 against the realized bits per weight of the trained allocation, for pretraining at 350M and 500M and for SFT.
Figure 15: Normalized sensitivity of Q-PACE and SNIP at budget 6.5 on {4,8,16} at 350M and 500M. Type shares over the calibrations, depth shares, and the final type-by-block map.
Figure 16: SFT training loss and gradient norm (mean of three seeds) and the normalized sensitivity by projection type of Q-PACE and SNIP at budget 6.5 on {4,8,16} .
Figure 17: 500M allocations of Q-PACE and SNIP at budget 6.5 on {4,8,16} over training. Parameter share per format and bits per weight by projection type.
Figure 18: Final 350M allocations of Q-PACE and SNIP for every budget and menu.
Figure 19: Final 500M allocations of Q-PACE and SNIP for every budget and menu.
Figure 20: Final SFT allocations of Q-PACE and SNIP for every budget and menu. Qwen3 has three MLP projections per block.
Figure 21: 4B normalized sensitivity at budget 6.5 on {4,16} . Type shares over the calibrations, depth shares, and the latest type-by-block map.
Figure 22: 4B sensitivity of all 168 layers across the calibrations at budget 6.5 on {4,16} . Colors are clipped at the 99th percentile, on separate scales for the two estimators.
Figure 23: 4B final allocations and their dynamics. Parameter share per format and bits per weight by projection type.
Figure 24: Final sensitivity (logarithmic color scale) and format of all 168 switchable layers in the four 4B Q-PACE runs.
run
GSM8K
IFEval
HumanEval
NQ
TQA
uniform references
uniform NVFP4
80.6±1.4
30.6±0.8
69.1±2.9
8.7±0.4
51.2±0.7
uniform MXFP8
84.2±0.3
34.3±1.1
75.8±0.9
11.2±0.7
52.6±0.1
budget 5, menu {4,6,8,16}
NVIDIA
81.0±0.9
28.0±1.7
71.3±2.7
9.2±0.5
51.2±0.2
Q-PACE
81.1±0.7
33.0±2.2
72.8±3.1
8.9±0.6
49.9±0.6
Appendix
Table 8: SFT generative and instruction benchmarks (accuracy %, mean ± seed standard deviation over three seeds).
run
ARC-c
ARC-e
HellaSw
PIQA
Wino
LAMBADA
MMLU
uniform references
uniform NVFP4
44.9±0.6
74.6±0.3
53.8±0.1
76.1±0.4
67.6±0.8
66.3±0.2
68.7±0.1
uniform MXFP8
47.0±0.2
78.4±0.1
54.6±0.0
78.0±0.1
70.7±0.6
68.6±0.1
71.1±0.1
budget 5, menu {4,6,8,16}
NVIDIA
44.6±0.5
74.6±0.7
53.7±0.2
76.1±0.3
67.9±1.2
67.0±0.2
68.8±0.1
Q-PACE
46.3±0.7
76.9±0.1
53.9±0.1
76.3±0.3
68.4±0.2
67.2±0.2
69.5±0.1
Appendix
Table 9: SFT knowledge/likelihood tasks (accuracy %, mean ± seed standard deviation over three seeds).
run
BoolQ
Copa
RACE
LogiQA
ASDiv
SAT-Math
NQ
uniform references
uniform NVFP4
77.1±0.2
84.0±1.7
39.5±0.7
40.2±0.7
27.9±0.2
43.6±2.0
8.7±0.4
uniform MXFP8
81.2±0.3
83.0±0.0
40.5±0.1
42.8±0.6
26.6±0.5
48.5±2.1
11.2±0.7
budget 5, menu {4,6,8,16}
NVIDIA
79.3±0.4
83.7±0.6
40.4±0.9
39.6±0.8
27.4±0.9
42.7±1.4
9.2±0.5
Q-PACE
81.9±0.6
83.7±3.1
40.1±0.9
41.0±0.8
28.5±1.0
45.8±1.3
8.9±0.6
Appendix
Table 10: SFT probe tasks (accuracy %, mean ± seed standard deviation over three seeds).