Large Vision-Language Models (LVLMs) frequently suffer from object hallucination. Existing training-free interventions primarily manipulate attention weights, which indirectly affect the deep semantics reaching the final predictive layers. In this work, we shift our focus to the hidden state vectors extracted after self-attention and residual addition. Empirical analysis reveals that hallucinated tokens do not simply over-rely on linguistic priors; instead, they exhibit an anomalous contextual deviation, showing significantly lower similarities to both textual and visual contexts in intermediate layers. Motivated by this, we propose ROT, a layer-specific, training-free framework. ROT dynamically detects semantic deviation in the middle layers and applies a norm-preserving rotation to steer the hidden states back toward the local multimodal context plane spanned by the contexts. For subsequent layers, a representational smoothing mechanism is introduced to stabilize the calibrated trajectory. Extensive experiments on multiple benchmarks demonstrate that ROT consistently reduces hallucinations across various model architectures and scales, offering an efficient, geometry-driven solution for grounded generation.
Figures & tables
Figure 1: The overall framework of the proposed ROT method. The process consists of two stages: (1) Dynamic Hallucination Detection in the critical middle layers, which identifies contextual deviation based on the similarities between the hidden state and multimodal contexts; and (2) Layer-Specific Hidden State Mitigation, which applies a norm-preserving rotation in the intermediate layers to steer the deviating representation back toward the grounded factual plane, followed by a representational smoothing mechanism in subsequent layers to maintain stability during propagation.
Figure 2: Average cosine similarities of the hidden state hk with the textual context ( ST ) and visual context ( SV ). Hallucinated tokens consistently exhibit significantly lower similarities to both contexts compared to normal tokens.
Figure 3: Distribution of the contextual deviation indicator E across different transformer layers for LLaVA-1.5-7B. The deviation is significantly more pronounced in the intermediate layers.
CHAIR
POPE
MME
Method
CS↓
CI↓
Recall ↑
length ↑
Acc ↑
F1 ↑
Exist. ↑
Pos. ↑
Color ↑
Total ↑
Greedy
54.4
15.7
75.2
82.7
85.5
85.9
175.67
114.00
151.00
440.67
VCD Leng et al. (2024)
51.1
14.7
75.2
82.3
85.0
85.3
184.67
128.67
153.00
466.34
OPERA Huang et al. (2024)
47.0
14.6
74.5
77.1
85.2
84.2
180.67
111.67
123.33
415.67
RITUAL Woo et al. (2024)
45.2
13.2
74.3
80.3
84.3
85.2
187.50
125.00
164.17
476.67
Vissink Kang et al. (2025)
52.4
14.5
75.1
83.4
86.5
86.0
190.00
138.33
155.00
483.33
Table 1: Performance comparison of ROT against various state-of-the-art training-free baselines on LLaVA-1.5-7B across CHAIR, POPE, and MME benchmarks.
CHAIR
POPE
Adversarial
Popular
Random
Method
CS↓
CI↓
Acc ↑
F1 ↑
Acc ↑
F1 ↑
Acc ↑
F1 ↑
Base Model: LLaVA-1.5-7B
Greedy
54.4
15.7
80.4
81.7
86.2
86.4
89.9
89.6
VCD
51.1
14.7
81.2
82.1
85.7
85.9
88.0
88.1
MemVR
50.4
14.3
79.8
81.8
86.3
85.9
89.4
89.6
Table 2: Generalization performance of ROT across various LVLM architectures and parameter scales on CHAIR and the POPE benchmark (including its adversarial, popular, and random subsets).
LLaVA-1.5-7B
Qwen2-VL-7B
Method
CS↓
CI↓
CS↓
CI↓
Greedy
54.5
15.7
24.2
8.0
ROT
33.0
6.7
21.4
4.9
w/o Rotation
40.5
12.4
23.4
6.6
w/o Smoothing
38.2
8.3
21.9
5.2
Table 3: Ablation study of ROT components on LLaVA-1.5-7B and Qwen2-VL-7B models across the CHAIR benchmark.
Figure 4: Qualitative comparison and hidden state similarities. The baseline model hallucinates a bench and a clock, which exhibit distinctly lower ST and SV compared to the factual token fence generated by ROT.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Qualitative Example 1. ROT successfully removes the hallucinated “bench” and “clock”, and accurately describes the “fence”.
Figure 6: Qualitative Example 2. ROT eliminates the hallucinated “handbag” and correctly identifies the “black and white” attribute of the image.
Figure 7: Qualitative Example 3. ROT intercepts the semantic deviation to prevent the hallucination of a “bench”.
Figure 8: Qualitative Example 4. ROT effectively suppresses clustered hallucinations (truck, traffic lights, handbags) and correctly recognizes the “bridge”.
Large Vision-Language Models (LVLMs) excel at multimodal tasks but remain prone to object hallucinations. Prior training-free remedies often uniformly strengthen visual signals, which may also amplify irrelevant regions and introduce spurious evidence, harming fluency. We propose Context-aware Attention Intervention (CAI), a training-free inference-time mechanism that enforces a see only when needed principle via two-axis selectivity: where to look and when to intervene. At each decoding step, CAI derives token-specific visual relevance from early-layer representations to localize semantically aligned regions, and applies a conservative, entropy- and depth-gated attention tilt only for uncertainty-spiking tokens in deeper layers where visual grounding degrades, leaving confident tokens and irrelevant regions largely unchanged. This targeted intervention strengthens visual grounding while preserving linguistic fluency, and it yields consistent improvements even without contrastive decoding, which remains optional as an auxiliary bias-suppression module. Extensive experiments across multiple LVLM backbones and benchmarks show that CAI achieves state-of-the-art hallucination mitigation, and our analysis characterizes CAI as a KL-minimal attention reweighting with bounded interference under inactive gates or small tilts. Code is available at https://github.com/Iris1946/CAI.
Yuqing Lei, Wenbo Lyu, Yingjun Du +3
University of Chinese Academy of Sciences · University of Amsterdam · United Imaging Healthcare Co., Ltd.
Large Vision-Language Models (LVLMs) have advanced multimodal understanding, yet their reliability is limited by hallucination, where generated content conflicts with visual facts. Existing mitigation methods either rely on costly external interventions, such as instruction tuning and retrieval, or use internal mechanisms that remain limited by flawed attention weights and entangled hidden representations. We propose Adversarial Orthogonal Disentanglement (AOD), a latent geometric framework for mitigating LVLM hallucinations. AOD learns a hallucination-related direction through a minimax objective: a classifier concentrates hallucination signals into the projected component, while an adversary removes them from the orthogonal residual space via a Gradient Reversal Layer. The learned direction enables a training-free dual-forward-pass contrastive decoding strategy that suppresses hallucinations while preserving general capabilities. Experiments on three LVLMs across four hallucination and four utility benchmarks show that AOD consistently outperforms strong baselines. It improves POPE accuracy by over 6% on average, boosts AMBER by 6%, and maintains strong performance on utility tasks such as MMMU. Further analysis shows robust transfer across datasets, suggesting that AOD captures general hallucination-related biases rather than dataset-specific artifacts. Our source code and datasets are available at https://github.com/Hunter-Wrynn/AOD.
Ruoxi Cheng, Haoxuan Ma, Zhengfei Hai +6
Fudan University · Tencent · Nanjing University +3
Large vision-language models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding. However, they remain prone to hallucinations, generating responses that are inconsistent with the visual evidence. Existing mitigation methods largely address language-prior bias or cross-modal imbalance, while progressive visual degradation across perception and memory remains underexplored. In this work, we propose Saliency-Driven Perceptual Realignment (SDPR), a training-free framework that mitigates the degradation of visual awareness throughout inference. Specifically, we first introduce saliency-driven attention redistribution to release attention hijacked by non-semantic sink tokens, thereby recovering critical visual evidence. Second, we identify spatial distortion in the KV cache and propose saliency-driven cache alignment to preserve query-relevant visual features during generation. Finally, we introduce prior-constrained contrastive decoding to penalize unfaithful predictions induced by dominant language priors. Our proposed SDPR is robust against hallucinations due to its holistic alignment of visual awareness across the entire generative trajectory. Extensive experiments across diverse LVLM architectures show that SDPR outperforms state-of-the-art methods on both hallucination and general-purpose benchmarks, requiring no additional training and incurring minimal runtime overhead. The code is available \href{https://github.com/PengSyuChen/SDPR}{\color{blue}{here}}.
Pengxu Chen, Yao Zhu, Guangming Zhu +4
Xidian University Xi’an, China · Tsinghua University Beijing, China · Shanghai Road Transport Development Center Shanghai, China +1