Post-training quantization of vision--language models (VLMs) is typically assessed through aggregate task accuracy and memory savings, but preserving a headline score does not guarantee preservation of visual grounding behavior. We present GHOST-Q, a cross-precision controlled evaluation of three 8B VLM families under FP16, INT8, and NF4 across utility and hallucination-sensitive benchmarks. Rather than comparing only aggregate accuracy, we pair FP16 and quantized predictions item by-item to quantify how compression redistributes grounding successes and failures. Five of six quantized variants preserve MMStar accuracy within ±2 percentage points, yet 10 of 36 paired effects remain significant after false-discovery-rate correction, nine on hallucination-sensitive conditions. Same-device A100 profiling further demonstrates that substantial memory reduction does not necessarily mean lower inference latency. Finally, an open-ended AMBER audit reveals strong generation budget censoring whose severity varies by architecture and precision. These results show that quantized VLMs should be evaluated jointly for aggregate utility, grounding reliability, generation behavior, and realized deployment efficiency.
Figures & tables
Figure 1: Overview of the Ghost-Q evaluation pipeline and main analyses.
Model
Prec.
MM*
POPE
A-E
A-A
A-R
Hall.
Macro
Qwen3
FP16
63.53
88.68
92.73
86.88
85.22
72.77
85.25
INT8
64.07
88.51
93.70
86.72
85.58
73.92
85.69
NF4
62.80
87.89
94.01
86.43
85.40
71.71
85.09
InternVL3
FP16
66.47
90.87
91.96
87.02
83.35
65.93
83.83
INT8
65.67
91.00
92.18
86.93
83.35
64.98
83.69
NF4
64.40
90.92
92.73
86.63
84.92
64.56
83.95
Table 1: Accuracy (%) across utility and grounding-sensitive tasks. “Macro” is the unweighted mean of POPE, three AMBER subsets, and HallusionBench.
Model
Task
Prec.
Δ
C → W
W → C
q
OR [95% CI]
Qwen3
A-exist
INT8
+0.97
4
52
<10−8
13.00 [4.78, 49.50]
Qwen3
A-exist
NF4
+1.28
14
77
<10−8
5.50 [3.09, 10.53]
Idefics3
POPE
NF4
-1.24
257
145
3.0×10−7
0.56 [0.46, 0.69]
Qwen3
POPE
NF4
-0.79
149
78
2.6×10−5
0.52 [0.39, 0.69]
Idefics3
A-exist
NF4
+1.22
67
127
1.4×10−4
1.90 [1.40, 2.59]
InternVL3
A-exist
NF4
+0.77
27
65
5.6×10−4
2.41 [1.52, 3.92]
Table 2: FDR-significant FP16–quantized paired effects. Δ is quantized minus FP16 accuracy (pp). C → W/W → C denote harmful/beneficial flips. q is the Benjamini–Hochberg-adjusted McNemar p -value. Matched OR >1 indicates more beneficial than harmful flips; OR <1 indicates the reverse.
Figure 2: Accuracy change relative to each architecture’s FP16 baseline. Stars mark comparisons significant after Benjamini–Hochberg correction over all 36 paired tests ( q<0.05 ).
Model
Prec.
Foot.
Peak
E2E 50
E2E 95
Thr.
(GiB)
(GiB)
(ms)
(ms)
ex/s
Qwen3
FP16
16.33
16.53
76.6
86.9
13.00
INT8
9.33
9.63
291.3
309.0
3.40
NF4
5.83
6.18
129.8
134.7
7.70
InternVL3
FP16
14.80
16.01
354.5
363.8
3.68
INT8
8.41
9.63
566.0
578.7
2.12
Table 3: Same-device A100 profile. E2E values are median/p95 latency over 600 measured inferences per configuration.
Figure 3: Measured footprint–latency trade-off on the same A100 GPU. Quantization consistently moves left (smaller footprint) but not down (faster inference); NF4 dominates INT8 in this two-dimensional efficiency plane for all three architectures.
Model
FP16
INT8
NF4
Qwen3-VL
794 (79.1%)
768 (76.5%)
734 (73.1%)
InternVL3
920 (91.6%)
913 (90.9%)
844 (84.1%)
Idefics3
1001 (99.7%)
1001 (99.7%)
1002 (99.8%)
Table 4: AMBER open-ended generation under a common 256-token budget. Entries are responses that hit the token ceiling out of 1,004.
Vision-Language Models (VLMs) achieve strong multimodal performance but are costly to deploy, and post-training quantization often causes significant accuracy loss. Despite its potential, quantization-aware training for VLMs remains underexplored. We propose GRACE, a framework unifying knowledge distillation and QAT under the Information Bottleneck principle: quantization constrains information capacity while distillation guides what to preserve within this budget. Treating the teacher as a proxy for task-relevant information, we introduce confidence-gated decoupled distillation to filter unreliable supervision, relational centered kernel alignment to transfer visual token structures, and an adaptive controller via Lagrangian relaxation to balance fidelity against capacity constraints. Across extensive benchmarks on LLaVA and Qwen families, our INT4 models consistently outperform FP16 baselines (e.g., LLaVA-1.5-7B: 70.1 vs. 66.8 on SQA; Qwen2-VL-2B: 76.9 vs. 72.6 on MMBench), nearly matching teacher performance. Using real INT4 kernel, we achieve 3× throughput with 54% memory reduction. This principled framework significantly outperforms existing quantization methods, making GRACE a compelling solution for resource-constrained deployment. Code and data are available at: https://github.com/ForeverBlue816/GRACE.
Department of Information Technology and Electrical Engineering, ETH Zurich, Zurich, Switzerland · School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore · Qualcomm AI Research, Amsterdam, the Netherlands +1
Post-training quantization reduces the memory requirements of vision-language-action (VLA) models, but precision selection must account for the interaction between layer scope, numerical format, and calibration. We introduce \textbf{VLAQuantBench}, a controlled evaluation with 409 runs and 94,574 simulation episodes: four models on LIBERO, with X-VLA additionally evaluated on three simulation benchmark families. Under uncalibrated W4A4 round-to-nearest quantization, expanding a π0.5 action-head subset from 126 to 167 layers raises success from 7.0% to 70.5%. Fixed-observation replay confirms a corresponding numerical recovery. Two-episode calibration removes the severe joint failures in the tested subsets, whereas the same smoothing-and-clipping recipe lowers π0 success and does not recover OpenVLA-OFT end-to-end. For OpenVLA-OFT, protecting one 28,672-parameter output projection instead restores near-baseline success: the remaining 441 eligible linear layers retain W3 on LIBERO-Long or eight-bit activations across all four suites. Task-clustered intervals support the large failure and recovery contrasts. These results establish recipe-dependent interactions and identify concrete precision assignments, rather than universal layer-sensitivity rules. Real-kernel and physical-robot measurements complement the accuracy analysis. Code, configurations, and episode records are publicly available at https://github.com/jiuyixu25/VLAQuantBench.
Jiuyi Xu, Qing Jin, Meida Chen +3
Colorado School of Mines · Independent Researcher · University of Central Florida +2
Recent inference-time hallucination mitigation methods for large vision-language models (LVLMs) report strong gains on hallucination benchmarks. However, it remains unclear whether lower hallucination scores reflect improved multimodal grounding or more conservative generation. We evaluate six mitigation methods across three LVLMs and four benchmarks, including hallucination-focused evaluation and the diverse capability benchmark MMStar. Our analysis reveals two consistent patterns. First, hallucination reduction is often coupled with reduced informativeness: methods that lower hallucination rates also reduce object recall, visual coverage, or response detailedness. Second, improvements on hallucination benchmarks do not reliably transfer to broader multimodal capabilities, with methods showing inconsistent or degraded performance on fine-grained perception and reasoning tasks. Our findings suggest that current evaluation protocols may overestimate progress by rewarding conservative generation. We argue that hallucination mitigation should be evaluated as a faithfulness--informativeness--capability trade-off rather than through hallucination scores alone.
Mehrdad Fazli, Sina Mansouri, Mohit Marvania +1
Department of Computer Science, George Mason University Fairfax, VA 22030, USA