cs.CVSep 24, 2026

Mind What Matters for Reasoning: Aligning Cross-Modal Attention via Selective Probability Mass Concentration

Authors: Jiaqi Deng, Zonghan Wu, Zhan Heng, Xiaoshui Huang, Huan Huo, Guandong Xu

Organizations: University of Technology Sydney · East China Normal University · The University of New South Wales · Shanghai Jiaotong University · The Education University of Hong Kong

Abstract

Multimodal large language models (MLLMs) achieve strong performance on visual reasoning tasks, yet remain prone to hallucinations and over-reliance on language priors, often generating answers without adequately using task-relevant visual evidence. Existing approaches primarily improve reasoning through reasoning-oriented supervision or inference-time strategies. In this work, we study a complementary question: can multimodal reasoning be improved by strengthening implicit visual grounding without directly supervising the reasoning process? Motivated by the functional specialization of attention heads, we investigate whether reasoning can be improved by guiding only the heads most responsive to visual evidence grounding. We propose Selective Probability Mass Concentration (sPMC), a training framework that identifies grounding-responsive heads and selectively regularizes their text-to-image attention. sPMC treats normalized attention over visual tokens as a spatial probability distribution and encourages the probability mass to be assigned to semantically relevant regions using segmentation-derived spatial priors. Adaptive Head Selection restricts this guidance to visually responsive heads while leaving the remaining heads unconstrained to preserve their complementary functions. Across 6 multimodal benchmark suites, sPMC achieves an average zero-shot improvement of 3% and gains of up to 11.3% across multiple MLLMs while regularizing only 3%-15% of their attention heads. These results demonstrate that targeted guidance of sparse and implicit visual evidence pathways can directly improve multimodal reasoning.

Figures & tables

Explore similar work

CardsList
  1. Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention

    May 21, 2026Changyuan Tian, Zhicong Lu, Huaxing Liu +7Multimodal ReasoningRecent Vision-Language Models

  2. Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs

    Aug 4, 2026Haoqian Kang, Liupeng Li, Kuofeng Gao +5Multimodal Large Language ModelsLLM Reasoning Strategies

  3. Look on Demand: A Cognitive Scheduling Framework for Visual Evidence Acquisition in Multimodal Reasoning

    May 27, 2026Yang Zhang, Xiaoshuai Sun, Rui Zhao +5Multimodal ReasoningMultimodal Reasoning Benchmarks