cs.CVSep 29, 2026

Beyond Attention Imbalance: Mitigating Hallucinations via Spectral Surgery

Authors: Siqi Lu, Suo Wei, Yongbin Zheng, Jianhang Yao, Wanying Xu, Peng Wang

Organizations: College of Intelligence Science and Technology, National University of Defense Technology, China · School of Computer Science, Northwestern Polytechnical University, China · Ningbo Institute, Northwestern Polytechnical University, China · National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean, China

Abstract

While Large Vision-Language Models (LVLMs) achieve remarkable success, hallucinations remain a significant barrier to their reliable deployment. Recent studies primarily attribute these issues to cross-modal attention imbalances; most solutions therefore focus on reweighting visual tokens or suppressing language priors. However, such approaches often overlook the spectral characteristics of the visual information flow and frequently rely on Contrastive Decoding (CD), which doubles inference time. Instead of following conventional approaches, we identify two distinct hallucination patterns-Perceptual-Semantic Dissociation and Localized Fixation-and propose FLASH (Frequency-Localized Attention SHaping), a training-free and CD-free framework. FLASH utilizes a Spectral Vortex Score to detect vision heads within multi-head attention layers and applies adaptive spectral modulation to rectify the visual information flow during decoding. Empirical results demonstrate that FLASH achieves a superior balance between performance and efficiency compared to SOTA methods.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Mitigating Hallucinations in Large Vision-Language Models via Causal Route Gating

    May 20, 2026Zhe Cheng, Wenyu Chen, Fode Zhang +1Recent Vision-Language ModelsHallucination Mitigation

  2. FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models

    Jun 28, 2026Yichen Guo, Kai Tang, Fenglai Lin +5Recent Vision-Language ModelsDecoding

  3. SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering

    Jul 5, 2026Kai Tang, Jinhao You, Bohua Zhang +6Large Language Model HallucinationObject Hallucination