cs.CVOct 4, 2026

Look Where You Say You're Looking: Self-Grounded Attention for Visual Reasoning

Authors: Uri Berger, Gal Chechik, Gal Dalal

Organizations: NVIDIA Research · The Hebrew University of Jerusalem · University of Melbourne · Bar-Ilan University

Abstract

We introduce Self-Saliency, a method for training Vision-Language Models (VLMs) to increase the alignment between their visual attention and the image regions mentioned in their reasoning. Self-Saliency uses a grounding model to localize the objects mentioned in each reasoning step and treats the resulting areas as supervision for the model's visual attention. Previous work on steering visual attention determines target image regions based solely on the image and question. In contrast, we show that conditioning the target regions on the model's generated reasoning improves downstream performance. For proper evaluation, we build a unified, broad suite of 25 visual reasoning benchmarks, where we reproduce the results of previous methods. We find that Self-Saliency significantly outperforms both prior attention-steering methods and baselines that ground image-level text, achieving both a better average rank and a better mean score. Post-training analysis shows that the model primarily adapts its reasoning text to existing attention patterns, producing shorter steps that refer to larger regions. Nevertheless, when controlling for generated text, attention to grounded regions increases significantly across the relevant layer. Finally, we identify a consistent geometric bias in VLM visual attention toward the image border. However, our ablations show that Self-Saliency's gains cannot be explained by simply aligning attention with the center of the image, highlighting the importance of aligning visual attention with the regions mentioned in the model's reasoning.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Attention-Guided Saliency Maps for Interpreting Visualization Literacy in VLMs

    Jul 17, 2026Maeve Hutchinson, Abderrahmane Wassim Mehdaoui, Pranava MadhyasthaSaliencyVisual Analytics

  2. ReGround: Restoring Visual Grounding in Multi-Step Reasoning through Self-Diagnosis and Visual Re-Examination

    Aug 5, 2026Lei Peng, Shuai Lv, Wei HuWeak Visual GroundingRecent Vision-Language Models

  3. Thinking with Visual Grounding

    Jun 15, 2026Junkai Zhang, Yihe Deng, Kai-Wei Chang +1Weak Visual GroundingVisual Reasoning