Look Where You Say You're Looking: Self-Grounded Attention for Visual Reasoning
Organizations: NVIDIA Research · The Hebrew University of Jerusalem · University of Melbourne · Bar-Ilan University
Abstract
We introduce Self-Saliency, a method for training Vision-Language Models (VLMs) to increase the alignment between their visual attention and the image regions mentioned in their reasoning. Self-Saliency uses a grounding model to localize the objects mentioned in each reasoning step and treats the resulting areas as supervision for the model's visual attention. Previous work on steering visual attention determines target image regions based solely on the image and question. In contrast, we show that conditioning the target regions on the model's generated reasoning improves downstream performance. For proper evaluation, we build a unified, broad suite of 25 visual reasoning benchmarks, where we reproduce the results of previous methods. We find that Self-Saliency significantly outperforms both prior attention-steering methods and baselines that ground image-level text, achieving both a better average rank and a better mean score. Post-training analysis shows that the model primarily adapts its reasoning text to existing attention patterns, producing shorter steps that refer to larger regions. Nevertheless, when controlling for generated text, attention to grounded regions increases significantly across the relevant layer. Finally, we identify a consistent geometric bias in VLM visual attention toward the image border. However, our ablations show that Self-Saliency's gains cannot be explained by simply aligning attention with the center of the image, highlighting the importance of aligning visual attention with the regions mentioned in the model's reasoning.
Figures & tables
| Type | Benchmarks |
| Mathematical reasoning | MathVista ( Lu et al., 2024 ) , MathVision ( Wang et al., 2024 ) , MathVerse ( Zhang et al., 2024 ) , WeMath ( Qiao et al., 2025 ) , MMK12 ( Meng et al., 2025 ) |
| Logical & algorithmic puzzles | LogicVista ( Xiao et al., 2024 ) , AlgoPuzzleVQA ( Ghosal et al., 2024 ) , VisuLogic ( Xu et al., 2026 ) , DailyClue ( Li et al., 2026b ) |
| Spatial & 3D reasoning | CV-Bench ( Zhu et al., 2025 ) , OmniSpatial ( Jia et al., 2026 ) |
| High-resolution visual search | V ∗ ( Wu and Xie, 2024 ) , HR-Bench 4K&8K ( Wang et al., 2025c ) , MME-RealWorld ( Zhang et al., 2025d ) |
| Hallucination & visual illusions | POPE ( Li et al., 2023 ) , HallusionBench ( Guan et al., 2024 ) , IllusionVQA ( Shahgir et al., 2024 ) |
| Fine-grained perception & chart reading | ChartQA ( Masry et al., 2022 ) , SalBench ( Huynh et al., 2025 ) |
| No attention intervention | Prior work | Ours | |||||
| Benchmark | Qwen3-VL 8B-Instruct | Coldstart | No-Sal | VGA | Saliency R1 | EASE | Self-Saliency |
| AlgoPuzzleVQA | 32.8 | 36.9 | 39.9 | 34.3 | 38.7 | 36.4 | 41.4 |
| ChartQA | 85.0 | 88.8 | 89.4 | 84.9 | 89.7 | 88.9 | 90.0 |
| CV-Bench | 86.0 | 83.4 | 84.9 | 86.1 | 83.8 | 84.2 | 85.9 |
| DailyClue | 35.4 | 32.4 | 34.2 | 34.2 | 32.0 | 32.7 | 31.2 |
| HallusionBench | 63.2 | 59.7 | 57.1 | 62.0 | 58.3 | 62.6 | 55.7 |
| Text-related | Attention-related | |||
| Completion | Box union | In-region share | In-region share | |
| length | area | (selected heads) | (layer-22 heads) | |
| Coldstart | 230.0 | 0.533 | 0.462 | 0.612 |
| Self-Saliency | 149.6 | 0.566 | 0.462 | 0.621 |
| Model | border ring | center | Largest |
| Qwen3-VL-8B-Instruct | 1.20 | 0.92 | Top left, 6.64 |
| InternVL3.5-8B | 1.18 | 0.95 | Bottom left, 2.29 |
| GLM-4.1V-Thinking | 1.50 | 0.82 | Top right, 7.79 |
| Nemotron-Omni | 1.43 | 0.87 | Top right, 3.86 |
| Human boxes | 0.53 | 1.14 | Non-border, 1.87 |
| Qwen3-VL-8B Instruct | coldstart | no_sal | center_rect | question_boxes | Self-Saliency | |
| Mean score ( ) | 61.14 | 62.69 | 63.47 | 63.39 | 63.33 | 64.26 |
| Mean rank ( ) | 3.90 | 4.06 | 3.34 | 3.68 | 3.38 | 2.64 |
| Wins no. ( ) | 7 | 0 | 3 | 2 | 5 | 9 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| No attention intervention | Prior work | Ours | |||||
| Benchmark | Qwen3-VL 8B-Instruct | Coldstart | No-Sal | VGA | Saliency- R1 | EASE | Self-Saliency |
| AlgoPuzzleVQA | 32.8 1.11 | 36.9 1.14 | 39.9 1.15 | 34.3 1.12 | 38.7 1.15 | 36.4 1.13 | 41.4 1.16 |
| ChartQA | 85.0 0.71 | 88.8 0.63 | 89.4 0.62 | 84.9 0.72 | 89.7 0.61 | 88.9 0.63 | 90.0 0.60 |
| CV-Bench | 86.0 0.68 | 83.4 0.72 | 84.9 0.70 | 86.1 0.67 | 83.8 0.72 | 84.2 0.71 | 85.9 0.68 |
| DailyClue | 35.4 1.85 | 32.4 1.82 | 34.2 1.84 | 34.2 1.84 | 32.0 1.81 | 32.7 1.82 | 31.2 1.80 |
| HallusionBench | 63.2 1.56 | 59.7 1.59 | 57.1 1.61 | 62.0 1.57 | 58.3 1.60 | 62.6 1.57 | 55.7 1.57 |
| No attention intervention | Saliency-based rewards | |||||
| Benchmark | Split | Qwen3-VL 8B-Instruct | Coldstart | No-Sal | Self-Saliency mean | Self-Saliency |
| AlgoPuzzleVQA | – | 32.8 | 36.9 | 39.9 | 38.8 | 41.4 |
| ChartQA | – | 85.0 | 88.8 | 89.4 | 89.7 | 90.0 |
| CV-Bench | – | 86.0 | 83.4 | 84.9 | 84.4 | 85.9 |
| DailyClue | – | 35.4 | 32.4 | 34.2 | 32.7 | 31.2 |
| HallusionBench | image | 63.2 | 59.7 | 57.1 | 55.7 | 55.7 |
| Model | Vision encoder | Grid shape | Language model |
| Qwen3-VL-8B-Instruct | ViT | Image-dependent | Qwen3-8B |
| InternVL3.5-8B | InternViT-300M | 16×16 | Qwen3-8B |
| GLM-4.1V-Thinking | AIMv2-Huge | Image-dependent | GLM-4-9B-0414 |
| Nemotron-Omni | C-RADIOv4-H | Image-dependent | Nemotron |