ROT: Rotating Hidden States towards Contextual Vectors for Hallucination Mitigation in LVLMs
Organizations: Harbin Institute of Technology, Shenzhen
Abstract
Large Vision-Language Models (LVLMs) frequently suffer from object hallucination. Existing training-free interventions primarily manipulate attention weights, which indirectly affect the deep semantics reaching the final predictive layers. In this work, we shift our focus to the hidden state vectors extracted after self-attention and residual addition. Empirical analysis reveals that hallucinated tokens do not simply over-rely on linguistic priors; instead, they exhibit an anomalous contextual deviation, showing significantly lower similarities to both textual and visual contexts in intermediate layers. Motivated by this, we propose ROT, a layer-specific, training-free framework. ROT dynamically detects semantic deviation in the middle layers and applies a norm-preserving rotation to steer the hidden states back toward the local multimodal context plane spanned by the contexts. For subsequent layers, a representational smoothing mechanism is introduced to stabilize the calibrated trajectory. Extensive experiments on multiple benchmarks demonstrate that ROT consistently reduces hallucinations across various model architectures and scales, offering an efficient, geometry-driven solution for grounded generation.
Figures & tables
| CHAIR | POPE | MME | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Method | Recall | length | Acc | F1 | Exist. | Pos. | Color | Total | ||
| Greedy | 54.4 | 15.7 | 75.2 | 82.7 | 85.5 | 85.9 | 175.67 | 114.00 | 151.00 | 440.67 |
| VCD Leng et al. (2024) | 51.1 | 14.7 | 75.2 | 82.3 | 85.0 | 85.3 | 184.67 | 128.67 | 153.00 | 466.34 |
| OPERA Huang et al. (2024) | 47.0 | 14.6 | 74.5 | 77.1 | 85.2 | 84.2 | 180.67 | 111.67 | 123.33 | 415.67 |
| RITUAL Woo et al. (2024) | 45.2 | 13.2 | 74.3 | 80.3 | 84.3 | 85.2 | 187.50 | 125.00 | 164.17 | 476.67 |
| Vissink Kang et al. (2025) | 52.4 | 14.5 | 75.1 | 83.4 | 86.5 | 86.0 | 190.00 | 138.33 | 155.00 | 483.33 |
| CHAIR | POPE | |||||||
| Adversarial | Popular | Random | ||||||
| Method | Acc | F1 | Acc | F1 | Acc | F1 | ||
| Base Model: LLaVA-1.5-7B | ||||||||
| Greedy | 54.4 | 15.7 | 80.4 | 81.7 | 86.2 | 86.4 | 89.9 | 89.6 |
| VCD | 51.1 | 14.7 | 81.2 | 82.1 | 85.7 | 85.9 | 88.0 | 88.1 |
| MemVR | 50.4 | 14.3 | 79.8 | 81.8 | 86.3 | 85.9 | 89.4 | 89.6 |
| LLaVA-1.5-7B | Qwen2-VL-7B | |||
|---|---|---|---|---|
| Method | ||||
| Greedy | 54.5 | 15.7 | 24.2 | 8.0 |
| ROT | 33.0 | 6.7 | 21.4 | 4.9 |
| w/o Rotation | 40.5 | 12.4 | 23.4 | 6.6 |
| w/o Smoothing | 38.2 | 8.3 | 21.9 | 5.2 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.