cs.CVApr 2, 2026

Attention at Rest Stays at Rest: Breaking Visual Inertia to Mitigate Relation Hallucinations

Authors: Boyang Gong, Yu Zheng, Fanye Kong, Jie Zhou, Jiwen Lu

Organizations: Tsinghua University

Abstract

While multimodal large language models demonstrate strong entity-level perception, faithfully grounding relational interactions between objects remains a persistent challenge. Although conventional visual grounding techniques attempt to resolve hallucinations by amplifying visual attention, strengthening overall visual signals fails to reliably correct relational errors. Tracing visual attention in relation descriptions reveals that correct responses tend to dynamically shift focus across regions, whereas hallucinated responses often linger on previously dominant evidence, exhibiting an undesirable \textit{visual inertia}. Further analysis shows that relation-prediction performance steadily deteriorates as more previous-step visual attention is carried into the current decoding step. We therefore introduce Inertia-aware Visual Excitation (IVE), an MLLM decoding method that dynamically recalibrates visual values using token-level attention history. By contrasting current attention against recent moving averages, IVE separates emergent tokens with rising relevance from persistently dominant inertia tokens, selectively reinforcing newly needed evidence while mildly attenuating contributions from repeatedly attended regions. Across three MLLMs and decoding strategies, IVE reduces relation hallucinations while preserving broader multimodal performance.

Explore similar work

CardsList
  1. When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs

    May 12, 2026Fanpu Cao, Xin Zou, Xuming Hu +1Large Language Model HallucinationMultimodal Large Language Models

  2. Correcting Visual Blur Induced by Attention Distraction to Reduce Hallucinations: Algorithm and Theory

    May 23, 2026Quanjiang Li, Zhiming Liu, Wei Luo +2Object HallucinationMultimodal Large Language Models

  3. Rethinking Visual Neglect: Steering via Context-Preference for MLLM Hallucination Mitigation

    May 27, 2026Jingwen Wu, Xijun Zhang, Ge SongHallucination MitigationMultimodal Large Language Models