Calibrate the Decisions That Change the Future: On-Policy Post-Training Quantization for Multimodal Large Language Models
Authors: Wenxiao Fan, Jingling Fu, Lichen Ma, Yu He, Luohang Liu, Jinbao Xue, Ke Zhang, Junshi Huang, +1 more
Organizations: School of Computer Science and Technology, Beijing Institute of Technology · Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University
Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences with local objectives. This overlooks autoregressive feedback: a quantization-induced token change redirects the prefix and changes future states. Yet on-policy coverage alone is insufficient because many decision mismatches barely affect future generation. We propose OnPTQ, an on-policy framework that calibrates on trajectories visited by the current quantized policy. On shared prefixes, OnPTQ identifies quantization-eroded boundaries, evaluates competing tokens through short counterfactual rollouts, and combines current discrepancy with branch consequence into a Decision--Consequence risk. The risk prioritizes critical states, while context anchoring and trajectory refresh preserve multimodal behavior and keep calibration aligned with the updated policy. We further derive a Decision--Consequence bound linking behavioral deviation to current policy discrepancy and action-conditioned future-value span. Across vision--language and omni-modal Qwen models under multiple low-bit settings, OnPTQ improves downstream performance and yields fewer correctness flips against the corresponding Dense/FP16 references, without changing the deployed inference graph.
Figures & tables
Figure 1: Motivating diagnostics for OnPTQ . (a) Quantization discrepancy amplifies after a decision fork. (b) On-policy calibration better covers deployment states. (c) A static anchor preserves downstream accuracy.
Figure 2: Overview of OnPTQ for a representative dense–quantized mismatch. Follow collects on-policy states, Compare identifies competing actions, Branch estimates future consequence, and Calibrate updates the quantizer before trajectory refresh.
Qwen2.5-VL-7B
Qwen3-VL-8B-Instruct
Method
Venue
Bits
MMMU
OCR
Viz
S-QA
T-VQA
Avg
MMMU
OCR
Viz
S-QA
T-VQA
Avg
Dense
–
FP16
46.7
83.8
70.8
88.4
82.9
74.5
51.6
86.2
69.3
94.6
81.6
76.6
RTN
–
W4A16
43.3
83.7
67.8
81.3
82.1
71.6
51.6
75.8
70.5
92.5
80.3
74.1
SmoothQuant
ICML’23
W4A16
40.7
79.4
67.3
81.9
81.7
70.2
49.0
71.5
70.0
93.1
79.9
72.7
MBQ
CVPR’25
W4A16
44.4
82.8
70.6
87.8
82.9
73.7
50.6
73.2
70.0
93.8
80.8
73.7
QIG
CVPR’26
W4A16
48.2
82.6
70.5
88.0
82.5
74.4
51.8
74.0
69.7
93.5
80.1
73.8
Table 1: Accuracy (%, ↑ ) across precision settings. OCR: OCRBench; Viz: VizWiz; S-QA: ScienceQA ; T-VQA: TextVQA. Avg: unweighted mean over the five benchmarks.
Qwen2.5-Omni-3B
Qwen2.5-Omni-7B
Audio–Text
Vision–Text
Omni-modal
Audio–Text
Vision–Text
Omni-modal
Method
Venue
Bits
Libri ↓
Wen ↓
MMMU ↑
Omni ↑
Libri ↓
Wen ↓
MMMU ↑
Omni ↑
Dense
–
FP16
3.9
7.5
43.3
43.8
2.9
7.1
50.0
45.3
RTN
–
W4A8
109.7
105.6
28.9
29.7
9.0
8.7
42.2
35.2
SmoothQuant
ICML’23
W4A8
77.4
94.2
30.0
27.3
8.6
8.3
42.2
36.7
MBQ
CVPR’25
W4A8
9.5
8.5
27.8
36.7
3.8
8.2
47.8
40.6
Table 2: Results on Qwen2.5-Omni under W4A8 and W4A6. Libri/Wen: WER (%, ↓ ) on LibriSpeech / WenetSpeech ; MMMU/Omni: accuracy (%, ↑ ) on MMMU/OmniBench.
Figure 3: Properties of Decision–Consequence risk. (a) High-consequence mismatches are sparse. (b) Mismatch weakly enriches future consequence. (c) Risk mass concentrates in few states. (d) Selected states migrate after one OnPTQ update.
Selection
Mean( et ) ↓
Pr(mQt<0)↓
Dense–InitQ
2.322
69.23
Mismatch-only
2.446
61.54
Mismatch + Near Miss
2.144
58.97
OnPTQ
1.609
53.85
Table 3: Fixed-budget selection on held-out joint-critical states.
Figure 4: Sample-wise Flip Rate relative to the same Dense/FP16 reference on Qwen3-VL-8B-Instruct (lower is better).
Table 5: End-to-end prefill performance of Qwen2.5-VL-7B on a single NVIDIA B200 (sequence length =2048 ). BS: batch size.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Theoretical object
OnPTQ instantiation
On-policy state st∼dQ
Quant-policy trajectory collection
Decision discrepancy
Boundary erosion et and same-state JS jt
Future consequence
Counterfactual branch amplification At
Policy-induced state shift
Iterative policy refresh
Appendix
Table 6: Correspondence between the Decision–Consequence analysis and the computable quantities used by OnPTQ .
Ord.
COCO image ID
Ord.
COCO image ID
Ord.
COCO image ID
Ord.
COCO image ID
0
000000039043
32
000000127647
64
000000222468
96
000000069410
1
000000153506
33
000000201420
65
000000044928
97
000000170040
2
000000098431
34
000000071044
66
000000117289
98
000000007232
3
000000130270
35
000000100034
67
000000235642
99
000000091912
4
000000016706
36
000000187474
68
000000211041
100
000000034785
5
000000173142
37
000000207597
69
000000066866
101
000000108531
Appendix
Table 7: COCO image IDs of the frozen 128-sample vision–language calibration manifest. Ord. denotes the zero-based manifest order.
Setting
MMMU
ScienceQA
Branch horizon H
2
47.9
88.9
4
48.8
89.7
8
48.3
89.2
Critical-state budget Kc
64
47.4
88.5
Appendix
Table 8: One-factor-at-a-time hyperparameter analysis on Qwen3-VL-8B-Instruct under W4A6. Shaded rows denote the main settings; all values are accuracy ( ↑ ).
Selection
All erosion
All flip (%)
Mismatch erosion
Mismatch flip (%)
Joint-critical erosion
Joint-critical flip (%)
Dense–InitQ
1.608
22.43
2.600
64.86
2.322
69.23
Mismatch-only
1.303
19.95
2.363
50.81
2.446
61.54
Mismatch + Near Miss
1.032
17.73
2.248
52.43
2.144
58.97
Consequence-only
1.073
15.91
2.003
45.95
1.664
43.59
OnPTQ
1.069
19.04
1.642
44.86
1.609
53.85
Random control
1.216
20.99
2.407
55.68
2.199
64.10
Appendix
Table 9: Complete fixed-state recovery results for alternative state-selection rules (lower is better).
Signal
Mismatch AUROC ↑
Future Spearman ↑
Hidden reconstruction error
0.470
-0.050
Same-state JS divergence
0.870
0.111
Boundary erosion et
0.821
0.077
Branch consequence At
0.543
0.504
OnPTQ risk rtDC
0.822
0.127
Appendix
Table 10: Diagnostics of individual signals and the OnPTQ risk for current mismatch and future consequence.