cs.CVOct 8, 2026

From Suppression to Repair: Mitigating Object Hallucination in Large Vision-Language Models via Localized Distribution Alignment

Authors: Chen Zhao, Xingping Dong, Jiachun Shi, Liang Peng, Chong Wang, Zhen Lei, Ran He, Bo Du

Organizations: School of Computer Science, Wuhan University · Institute of Automation, Chinese Academy of Sciences

Abstract

Object hallucination remains a major obstacle for large vision-language models (LVLMs) to generate reliable content. An intuitive mitigation strategy is to suppress hallucination-related components in hidden representations. However, these components may also contain useful information, and suppressing them can weaken the model's multimodal capabilities. In this paper, we propose ResOT, a training-free method that repairs representations at inference time through localized distribution alignment. Specifically, ResOT projects dominant hallucinated directions away from the faithful subspace, forming a low-dimensional residual subspace for intervention. Within this subspace, ResOT uses Gaussian optimal transport (OT) to align the hallucinated distribution with the faithful one. The resulting map defines repair targets with minimal changes to the original representations. At inference, ResOT adaptively controls how far each token state moves toward its OT target. Experiments on three representative LVLMs show that ResOT substantially reduces object hallucination while improving image caption quality and multimodal performance across multiple benchmarks. Code will be released.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Mitigating Object Hallucinations in Vision-Language Models through Region-Aware Attention Recalibration

    May 24, 2026Yuanzhi Xu, Qian Gao, Jun Fan +4VLM RobustnessVision-Language Models

  2. Test-Time Hallucination Control in Large Vision-Language Models

    Aug 11, 2026Mehran Tamjidi, Hamidreza Dastmalchi, Ali Cheraghian +3Test-Time Adaptation of VLMsLarge Vision-Language Models

  3. Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation

    May 25, 2026Ruoxi Cheng, Haoxuan Ma, Zhengfei Hai +6Adversarial Representation LearningVision-Language Models