cs.AISep 29, 2026

Calibrate the Decisions That Change the Future: On-Policy Post-Training Quantization for Multimodal Large Language Models

Authors: Wenxiao Fan, Jingling Fu, Lichen Ma, Yu He, Luohang Liu, Jinbao Xue, Ke Zhang, Junshi Huang, +1 more

Organizations: School of Computer Science and Technology, Beijing Institute of Technology · Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University

Abstract

Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences with local objectives. This overlooks autoregressive feedback: a quantization-induced token change redirects the prefix and changes future states. Yet on-policy coverage alone is insufficient because many decision mismatches barely affect future generation. We propose OnPTQ, an on-policy framework that calibrates on trajectories visited by the current quantized policy. On shared prefixes, OnPTQ identifies quantization-eroded boundaries, evaluates competing tokens through short counterfactual rollouts, and combines current discrepancy with branch consequence into a Decision--Consequence risk. The risk prioritizes critical states, while context anchoring and trajectory refresh preserve multimodal behavior and keep calibration aligned with the updated policy. We further derive a Decision--Consequence bound linking behavioral deviation to current policy discrepancy and action-conditioned future-value span. Across vision--language and omni-modal Qwen models under multiple low-bit settings, OnPTQ improves downstream performance and yields fewer correctness flips against the corresponding Dense/FP16 references, without changing the deployed inference graph.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Reconstruction Loss in Post-Training Quantization: Balanced Fitting for Large Vision-Language Models

    Sep 28, 2026Minchan Kang, Kyeonghye Park, Seungyeon Sa +3Post-Training QuantizationRecent Vision-Language Models

  2. Coverage-Based Calibration for Post-Training Quantization via Weighted Set Cover over Outlier Channels

    Apr 27, 2026Ibne Farabi Shihab, Sanjeda Akter, Anuj SharmaPost-Training QuantizationPost-Training