Calibrate the Decisions That Change the Future: On-Policy Post-Training Quantization for Multimodal Large Language Models
Organizations: School of Computer Science and Technology, Beijing Institute of Technology · Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University
Abstract
Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences with local objectives. This overlooks autoregressive feedback: a quantization-induced token change redirects the prefix and changes future states. Yet on-policy coverage alone is insufficient because many decision mismatches barely affect future generation. We propose OnPTQ, an on-policy framework that calibrates on trajectories visited by the current quantized policy. On shared prefixes, OnPTQ identifies quantization-eroded boundaries, evaluates competing tokens through short counterfactual rollouts, and combines current discrepancy with branch consequence into a Decision--Consequence risk. The risk prioritizes critical states, while context anchoring and trajectory refresh preserve multimodal behavior and keep calibration aligned with the updated policy. We further derive a Decision--Consequence bound linking behavioral deviation to current policy discrepancy and action-conditioned future-value span. Across vision--language and omni-modal Qwen models under multiple low-bit settings, OnPTQ improves downstream performance and yields fewer correctness flips against the corresponding Dense/FP16 references, without changing the deployed inference graph.
Figures & tables
| Qwen2.5-VL-7B | Qwen3-VL-8B-Instruct | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | Venue | Bits | MMMU | OCR | Viz | S-QA | T-VQA | Avg | MMMU | OCR | Viz | S-QA | T-VQA | Avg |
| Dense | – | FP16 | 46.7 | 83.8 | 70.8 | 88.4 | 82.9 | 74.5 | 51.6 | 86.2 | 69.3 | 94.6 | 81.6 | 76.6 |
| RTN | – | W4A16 | 43.3 | 83.7 | 67.8 | 81.3 | 82.1 | 71.6 | 51.6 | 75.8 | 70.5 | 92.5 | 80.3 | 74.1 |
| SmoothQuant | ICML’23 | W4A16 | 40.7 | 79.4 | 67.3 | 81.9 | 81.7 | 70.2 | 49.0 | 71.5 | 70.0 | 93.1 | 79.9 | 72.7 |
| MBQ | CVPR’25 | W4A16 | 44.4 | 82.8 | 70.6 | 87.8 | 82.9 | 73.7 | 50.6 | 73.2 | 70.0 | 93.8 | 80.8 | 73.7 |
| QIG | CVPR’26 | W4A16 | 48.2 | 82.6 | 70.5 | 88.0 | 82.5 | 74.4 | 51.8 | 74.0 | 69.7 | 93.5 | 80.1 | 73.8 |
| Qwen2.5-Omni-3B | Qwen2.5-Omni-7B | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Audio–Text | Vision–Text | Omni-modal | Audio–Text | Vision–Text | Omni-modal | |||||
| Method | Venue | Bits | Libri | Wen | MMMU | Omni | Libri | Wen | MMMU | Omni |
| Dense | – | FP16 | 3.9 | 7.5 | 43.3 | 43.8 | 2.9 | 7.1 | 50.0 | 45.3 |
| RTN | – | W4A8 | 109.7 | 105.6 | 28.9 | 29.7 | 9.0 | 8.7 | 42.2 | 35.2 |
| SmoothQuant | ICML’23 | W4A8 | 77.4 | 94.2 | 30.0 | 27.3 | 8.6 | 8.3 | 42.2 | 36.7 |
| MBQ | CVPR’25 | W4A8 | 9.5 | 8.5 | 27.8 | 36.7 | 3.8 | 8.2 | 47.8 | 40.6 |
| Selection | Mean( ) | |
|---|---|---|
| Dense–InitQ | 2.322 | 69.23 |
| Mismatch-only | 2.446 | 61.54 |
| Mismatch + Near Miss | 2.144 | 58.97 |
| OnPTQ | 1.609 | 53.85 |
| Setting | MMMU | ScienceQA |
|---|---|---|
| State Selection | ||
| All on-policy states | 45.4 | 86.8 |
| Mismatch-only | 47.1 | 88.3 |
| Risk Composition | ||
| Boundary only: | 44.9 | 86.5 |
| w/o consequence | 45.7 | 87.3 |
| Metric | BS | FP16 | W4A8 |
|---|---|---|---|
| Peak Memory (GB) | 1 | 16.22 | 10.15 |
| Prefill (ms) | 1 | 36.73 | 20.48 |
| Peak Memory (GB) | 8 | 21.13 | 15.06 |
| Prefill (ms) | 8 | 271.39 | 192.76 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Theoretical object | OnPTQ instantiation |
|---|---|
| On-policy state | Quant-policy trajectory collection |
| Decision discrepancy | Boundary erosion and same-state JS |
| Future consequence | Counterfactual branch amplification |
| Policy-induced state shift | Iterative policy refresh |
| Ord. | COCO image ID | Ord. | COCO image ID | Ord. | COCO image ID | Ord. | COCO image ID |
|---|---|---|---|---|---|---|---|
| 0 | 000000039043 | 32 | 000000127647 | 64 | 000000222468 | 96 | 000000069410 |
| 1 | 000000153506 | 33 | 000000201420 | 65 | 000000044928 | 97 | 000000170040 |
| 2 | 000000098431 | 34 | 000000071044 | 66 | 000000117289 | 98 | 000000007232 |
| 3 | 000000130270 | 35 | 000000100034 | 67 | 000000235642 | 99 | 000000091912 |
| 4 | 000000016706 | 36 | 000000187474 | 68 | 000000211041 | 100 | 000000034785 |
| 5 | 000000173142 | 37 | 000000207597 | 69 | 000000066866 | 101 | 000000108531 |
| Setting | MMMU | ScienceQA |
|---|---|---|
| Branch horizon | ||
| 2 | 47.9 | 88.9 |
| 4 | 48.8 | 89.7 |
| 8 | 48.3 | 89.2 |
| Critical-state budget | ||
| 64 | 47.4 | 88.5 |
| Selection | All erosion | All flip (%) | Mismatch erosion | Mismatch flip (%) | Joint-critical erosion | Joint-critical flip (%) |
|---|---|---|---|---|---|---|
| Dense–InitQ | 1.608 | 22.43 | 2.600 | 64.86 | 2.322 | 69.23 |
| Mismatch-only | 1.303 | 19.95 | 2.363 | 50.81 | 2.446 | 61.54 |
| Mismatch + Near Miss | 1.032 | 17.73 | 2.248 | 52.43 | 2.144 | 58.97 |
| Consequence-only | 1.073 | 15.91 | 2.003 | 45.95 | 1.664 | 43.59 |
| OnPTQ | 1.069 | 19.04 | 1.642 | 44.86 | 1.609 | 53.85 |
| Random control | 1.216 | 20.99 | 2.407 | 55.68 | 2.199 | 64.10 |
| Signal | Mismatch AUROC | Future Spearman |
|---|---|---|
| Hidden reconstruction error | 0.470 | -0.050 |
| Same-state JS divergence | 0.870 | 0.111 |
| Boundary erosion | 0.821 | 0.077 |
| Branch consequence | 0.543 | 0.504 |
| OnPTQ risk | 0.822 | 0.127 |