Object hallucination remains a major obstacle for large vision-language models (LVLMs) to generate reliable content. An intuitive mitigation strategy is to suppress hallucination-related components in hidden representations. However, these components may also contain useful information, and suppressing them can weaken the model's multimodal capabilities. In this paper, we propose ResOT, a training-free method that repairs representations at inference time through localized distribution alignment. Specifically, ResOT projects dominant hallucinated directions away from the faithful subspace, forming a low-dimensional residual subspace for intervention. Within this subspace, ResOT uses Gaussian optimal transport (OT) to align the hallucinated distribution with the faithful one. The resulting map defines repair targets with minimal changes to the original representations. At inference, ResOT adaptively controls how far each token state moves toward its OT target. Experiments on three representative LVLMs show that ResOT substantially reduces object hallucination while improving image caption quality and multimodal performance across multiple benchmarks. Code will be released.
Figures & tables
Figure 1: Conceptual comparison of suppression and repair. (a) Faithful (blue) and hallucinated (orange) representations form partially overlapping clusters. (b) Suppression compresses the orange cluster into a narrow band and also displaces some faithful points (blue arrows). (c) Repair instead moves the orange points toward the blue cluster (pink arrows), forming the repaired cluster in pink.
Figure 2: Overview of ResOT . Where to repair (top) : the orange hallucinated directions are projected away from the blue faithful subspace to form the green residual subspace. How to repair (bottom): Gaussian OT aligns the orange and blue residual distributions to define repair targets (left); a likelihood-based gate controls the movement toward each target during inference (right).
Method
LLaVA-1.5
mPLUG-Owl2
InstructBLIP
Acc. ↑
Prec. ↑
F1 ↑
Acc. ↑
Prec. ↑
F1 ↑
Acc. ↑
Prec. ↑
F1 ↑
Vanilla
84.20
82.30
84.67
78.23
72.07
81.07
84.17
83.07
84.60
OPERA
84.67
83.37
85.13
79.77
74.37
82.03
85.00
83.93
85.43
VCD
81.13
77.83
82.30
76.73
70.83
79.77
75.97
78.20
81.63
HALC
84.83
82.80
85.53
78.30
72.23
81.13
84.60
83.37
85.13
VISTA
84.23
80.03
85.47
76.00
69.33
79.70
85.20
84.87
85.50
Table 1: Average POPE results across the random, popular and adversarial splits. Higher values indicate better performance. The best results are highlighted in bold.
Method
LLaVA-1.5
mPLUG-Owl2
InstructBLIP
Cs↓
Ci↓
BLEU-1 ↑
Cs↓
Ci↓
BLEU-1 ↑
Cs↓
Ci↓
BLEU-1 ↑
Vanilla
0.514
0.1451
0.1739
0.568
0.1639
0.1559
0.462
0.1310
0.1740
OPERA
0.454
0.1317
0.1769
0.488
0.1503
0.1769
0.398
0.1242
0.1524
VCD
0.484
0.1437
0.1743
0.580
0.1808
0.1689
0.566
0.1808
0.1698
HALC
0.380
0.1270
0.1870
0.500
0.1432
0.1762
0.620
0.1770
0.1580
VISTA
0.436
0.1336
0.1619
0.528
0.1638
0.1426
0.434
0.1345
0.1508
Table 2: Captioning results on COCO evaluated by CHAIR and BLEU-1. Lower Cs and Ci indicate fewer hallucinated objects, while higher BLEU-1 indicates better caption quality. The maximum generation length is set to 512 tokens. The best results are highlighted in bold.
Backbone
Vanilla vs. ResOT
Nullu vs. ResOT
Acc.
Det.
Acc.
Det.
LLaVA-1.5
5.72 / 6.77
6.09 / 5.81
6.22 / 6.32
5.76 / 5.82
mPLUG-Owl2
5.65 / 6.52
5.90 / 5.94
6.38 / 6.88
6.22 / 6.24
InstructBLIP
5.68 / 6.38
5.82 / 5.94
5.36 / 6.48
5.72 / 6.00
Table 3: GPT-4o evaluation on COCO image descriptions. Each entry reports the baseline/ResOT scores for visual accuracy (Acc.) and detailedness (Det.); higher values indicate better performance.
Figure 3: MME results across three LVLM backbones. Each cell reports the raw score, while color intensity indicates the normalized score within each metric. We report the total, perception and cognition scores, together with four perception subcategories: existence, color, position and count.
Figure 4: Representation visualization at layer 24. From left to right: the original representation space, the Nullu-style subspace, and the ResOT residual subspace before and after repair. Faithful and hallucinated representations are evaluated on the held-out split. We use PCA for visualization.
Variant
Cs↓
BLEU-1 ↑
MME ↑
RandSub + OT
0.516
0.1748
1787.17
NulluSub + OT
0.424
0.1681
1764.82
ResSub + Sup
0.378
0.1903
1784.96
ResSub + Mean
0.320
0.2140
1798.35
ResSub + OT w/o Gate
0.292
0.2268
1814.17
ResSub + OT w/ Gate
0.292
0.2334
1819.02
Table 4: Ablation study results on subspace construction and intervention operations.
Figure 5: Effect of calibration size and corruption. We compare ResOT and Nullu under different calibration sizes and corruption rates. Error bars denote the standard deviation across random trials.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Method
BLEU-1
BLEU-2
BLEU-3
BLEU-4
ROUGE-L
CLIPScore
SPICE
LLaVA-1.5
Vanilla
0.1739
0.1174
0.0738
0.0465
0.1804
0.8174
0.1827
Nullu
0.1581
0.1091
0.0689
0.0429
0.1707
0.8215
0.1761
ResOT
0.2334
0.1554
0.0979
0.0609
0.2062
0.8251
0.1961
mPLUG-Owl2
Vanilla
0.1559
0.1157
0.0740
0.0472
0.1788
0.8180
0.1783
Nullu
0.1830
0.1243
0.0788
0.0496
0.1915
0.8217
0.1926
ResOT
0.2293
0.1565
0.1014
0.0654
0.2213
0.8281
0.2105
Appendix
Table 5: Comparison of caption quality across different LVLMs. Higher values are better.
Backbone
Method
CHAIR ↓
Cover ↑
Hal ↓
Cog ↓
LLaVA-1.5
Vanilla
7.6
48.7
35.4
4.2
Nullu
7.6
49.2
37.8
3.7
ResOT
6.0
48.2
25.1
2.4
mPLUG-Owl2
Vanilla
9.3
50.3
40.8
5.3
Nullu
6.6
48.3
27.7
3.0
ResOT
5.0
48.9
21.3
1.8
Appendix
Table 6: AMBER results across three LVLM backbones. Lower CHAIR, Hal, and Cog indicate less hallucination, while higher Cover indicates better object coverage.
Method
LLaVA-1.5
mPLUG-Owl2
InstructBLIP
Avg. ↑
Hall. ↓
Avg. ↑
Hall. ↓
Avg. ↑
Hall. ↓
Vanilla
2.1250
0.6458
1.9479
0.6667
1.8229
0.6458
Nullu
2.1146
0.6562
1.9583
0.6458
1.6250
0.6979
ResOT
2.2708
0.5938
2.0104
0.6458
1.9583
0.6146
Appendix
Table 7: MMHal-Bench results across three LVLM backbones. Avg. denotes average score, and Hall. denotes hallucination rate.
Backbone
Method
Rec.
OCR
Know.
Gen.
Spat.
Math
Total
LLaVA-1.5
Vanilla
34.0
22.6
16.1
20.1
24.4
11.5
29.5
ResOT
34.6
23.2
16.3
19.4
27.1
11.5
31.0
mPLUG-Owl2
Vanilla
38.7
27.0
23.8
25.4
29.5
3.8
34.3
ResOT
38.2
28.2
23.6
24.9
30.9
3.8
34.5
InstructBLIP
Vanilla
30.9
15.9
15.4
16.0
17.3
8.1
25.8
ResOT
33.5
17.1
19.6
18.0
18.1
7.3
27.7
Appendix
Table 8: MM-Vet results across three LVLM backbones. We report scores on recognition (Rec.), OCR, knowledge (Know.), generation (Gen.), spatial reasoning (Spat.), math, and the overall total score. Higher scores indicate better performance.
Model
Method
MMMU
TextVQA
ScienceQA
MMBench
LLaVA-1.5
Vanilla
35.22
45.87
65.25
61.76
Nullu
34.89
43.04
64.95
59.29
ResOT
35.89
46.15
65.49
61.76
mPLUG-Owl2
Vanilla
36.56
55.36
68.07
63.62
Nullu
36.11
55.04
68.02
63.24
ResOT
36.67
55.17
68.22
63.85
Appendix
Table 9: Results on MMMU, TextVQA, ScienceQA, and MMBench across three LVLM backbones. Higher values are better.
Method
Tokens/s (vs. Vanilla) ↑
Peak Memory Increase (MiB) ↓
Preserves LVLM Weights
OPERA
0.752 ( −97.87 %)
25219.631
✓
VCD
17.462 ( −52.64 %)
724.160
✓
HALC
1.858 ( −95.21 %)
4728.521
✓
ICT
24.285 ( −39.46 %)
0.250
✓
VTI
29.940 ( −24.77 %)
56.504
✓
VISTA
22.200 ( −51.43 %)
3761.718
✓
Appendix
Table 10: Inference efficiency on LLaVA-1.5 with 128 generated tokens. Changes are computed against each method’s paired Vanilla run.
Figure 6: Representation visualizations across layers. Rows correspond to layers 16, 24, and 31. Columns show the original representation space, the Nullu-style subspace, and the ResOT residual subspace before and after repair. We use PCA for visualization.
Figure 7: Sensitivity analysis of the repair strength α and residual subspace rank on LLaVA-1.5. For α sensitivity, we fix the rank to 32. For rank sensitivity, we fix α to 0.7.
Layer
Setting
Mahalanobis Dist. ↑
LDA Acc. ↑
LDA AUC ↑
16
Nullu-style
1.5033
0.7795
0.8674
ResOT before repair
2.4725
0.8865
0.9544
24
Nullu-style
1.3405
0.7590
0.8366
ResOT before repair
2.1708
0.8565
0.9351
31
Nullu-style
1.2530
0.7325
0.8161
ResOT before repair
1.9297
0.8460
0.9149
Appendix
Table 11: Separability between faithful and hallucinated LLaVA-1.5 representations in the subspace on the evaluation split. Higher values indicate stronger separability.
Layer
Nullu ( F→H )
Before repair ( F→H )
After repair ( F→R )
16
1.9758
1.2845
0.0537
24
4.0293
2.9075
0.1265
31
4.4187
2.5757
0.0994
Appendix
Table 12: Center-distance analysis across layers. F , H , and R denote faithful, hallucinated, and repaired representations, respectively.
Layer
Faithful (%)
Hallucinated (%)
H/F
16
1.5903
7.1353
4.49 ×
24
1.1953
5.9745
5.00 ×
31
1.3840
4.2196
3.05 ×
16–31
1.3425
6.0281
4.49 ×
Appendix
Table 13: Average relative displacement of faithful and hallucinated samples in the representation space. H/F denotes the ratio of hallucinated to faithful displacement. The 16–31 row reports the average across all selected layers.
Figure 8: Gaussian diagnostics at layer 31. Columns show histograms of whitened coordinates (left), normal Q–Q plots (middle), and Mahalanobis Q–Q plots (right). Green and red denote faithful and hallucinated representations, respectively. Dashed black lines show the standard normal reference in the histograms and y=x in the Q–Q plots.
Figure 9: Gaussian diagnostics across layers 16–31. Columns report the KS distance of whitened coordinate values to N(0,1) (left) and the KS distance of squared Mahalanobis distances to χk2 (right). Green and red curves denote faithful and hallucinated representations, respectively.
The generation of factually incorrect objects, commonly known as object hallucination, remains a persistent challenge in Large Vision-Language Models (LVLMs). Current approaches to address this issue - ranging from expensive data-driven fine-tuning and high-latency contrastive decoding to rigid attention head truncation - frequently compromise either computational efficiency or the continuity of the model's feature space. To overcome these limitations, we introduce a novel, training-free inference strategy that operates as a region-aware adaptive weighting mechanism to dynamically correct semantic drift without relying on abrupt heuristic truncations. By computing an outlier-resistant statistical midpoint across various attention heads, we establish a stable anchor for reliable visual representations. We then utilize the inter-head disagreement mapped across regions to dynamically determine intervention budgets, gently suppressing hallucination-inducing attention paths through a continuous penalty modulation. This recalibration process effectively rectifies visual-semantic misalignments while fully preserving generative fluency and language priors. Comprehensive evaluations on standard multimodal benchmarks, including CHAIR, POPE, and MME, reveal that our strategy substantially curtails both instance- and sentence-level hallucinations. The results demonstrate state-of-the-art performance against contemporary baselines, confirming our method's efficiency and algorithmic robustness. Our code will be public.
Yuanzhi Xu, Qian Gao, Jun Fan +4
Qilu University of Technology (Shandong Academy of Sciences) Jinan, China · China Telecom Digital Intelligence Technology Co, Ltd Jinan, China · Shenyang Aerospace University Shenyang, China +1
Object Hallucination in large vision-language models (LVLMs), where models generate non-factual content about input images, remains a critical barrier to their reliability in real-world applications. Existing mitigation strategies can be categorized into training-based and training-free methods. Training-based methods often achieve strong performance but are costly, requiring extensive computational resources, large-scale data, and time-consuming fine-tuning. Training-free approaches are particularly appealing due to their efficiency. However, existing training-free methods either require multiple decoding rounds, which adds computational overhead, or modify internal states in a model-specific way that risks degrading pretrained knowledge. We propose Test-Time Hallucination Mitigation (TTH) method, a novel training-free method that addresses both limitations. TTH introduces a token-validator module, implemented as a zero-shot Multi-Modal Classifier (MMC), to generate auxiliary logits grounded in the input image. These logits are fused with the original LVLM outputs at the token level for object tokens selected from a candidate pool. An entropy-based weighting scheme is then applied to enable robust and accurate predictions. Extensive experiments across multiple LVLM families and diverse benchmarks demonstrate that TTH consistently improves accuracy and robustness, underscoring its generalizability and practical effectiveness. Code is released at https://github.com/Mehran-TAM/TTH
Mehran Tamjidi, Hamidreza Dastmalchi, Ali Cheraghian +3
University of Technology Sydney · York University, Canada · Australian National University +2
Large Vision-Language Models (LVLMs) have advanced multimodal understanding, yet their reliability is limited by hallucination, where generated content conflicts with visual facts. Existing mitigation methods either rely on costly external interventions, such as instruction tuning and retrieval, or use internal mechanisms that remain limited by flawed attention weights and entangled hidden representations. We propose Adversarial Orthogonal Disentanglement (AOD), a latent geometric framework for mitigating LVLM hallucinations. AOD learns a hallucination-related direction through a minimax objective: a classifier concentrates hallucination signals into the projected component, while an adversary removes them from the orthogonal residual space via a Gradient Reversal Layer. The learned direction enables a training-free dual-forward-pass contrastive decoding strategy that suppresses hallucinations while preserving general capabilities. Experiments on three LVLMs across four hallucination and four utility benchmarks show that AOD consistently outperforms strong baselines. It improves POPE accuracy by over 6% on average, boosts AMBER by 6%, and maintains strong performance on utility tasks such as MMMU. Further analysis shows robust transfer across datasets, suggesting that AOD captures general hallucination-related biases rather than dataset-specific artifacts. Our source code and datasets are available at https://github.com/Hunter-Wrynn/AOD.
Ruoxi Cheng, Haoxuan Ma, Zhengfei Hai +6
Fudan University · Tencent · Nanjing University +3