We introduce Self-Saliency, a method for training Vision-Language Models (VLMs) to increase the alignment between their visual attention and the image regions mentioned in their reasoning. Self-Saliency uses a grounding model to localize the objects mentioned in each reasoning step and treats the resulting areas as supervision for the model's visual attention. Previous work on steering visual attention determines target image regions based solely on the image and question. In contrast, we show that conditioning the target regions on the model's generated reasoning improves downstream performance. For proper evaluation, we build a unified, broad suite of 25 visual reasoning benchmarks, where we reproduce the results of previous methods. We find that Self-Saliency significantly outperforms both prior attention-steering methods and baselines that ground image-level text, achieving both a better average rank and a better mean score. Post-training analysis shows that the model primarily adapts its reasoning text to existing attention patterns, producing shorter steps that refer to larger regions. Nevertheless, when controlling for generated text, attention to grounded regions increases significantly across the relevant layer. Finally, we identify a consistent geometric bias in VLM visual attention toward the image border. However, our ablations show that Self-Saliency's gains cannot be explained by simply aligning attention with the center of the image, highlighting the importance of aligning visual attention with the regions mentioned in the model's reasoning.
Figures & tables
Figure 1: Self-Saliency is trained to increase the alignment between visual attention and the regions referenced in its reasoning chain, creating two complementary pressures: (1) generate claims that can be grounded in the image, and (2) attend to the regions those claims refer to.
Figure 2: Our saliency-scoring pipeline. Given an input image and reasoning step, we extract bounding boxes corresponding to the mentioned image regions, compute a saliency map from the model’s visual attention, and calculate the saliency score based on the overlap between the bounding boxes and the saliency map.
Type
Benchmarks
Mathematical reasoning
MathVista ( Lu et al., 2024 ) , MathVision ( Wang et al., 2024 ) , MathVerse ( Zhang et al., 2024 ) , WeMath ( Qiao et al., 2025 ) , MMK12 ( Meng et al., 2025 )
Logical & algorithmic puzzles
LogicVista ( Xiao et al., 2024 ) , AlgoPuzzleVQA ( Ghosal et al., 2024 ) , VisuLogic ( Xu et al., 2026 ) , DailyClue ( Li et al., 2026b )
Spatial & 3D reasoning
CV-Bench ( Zhu et al., 2025 ) , OmniSpatial ( Jia et al., 2026 )
High-resolution visual search
V ∗ ( Wu and Xie, 2024 ) , HR-Bench 4K&8K ( Wang et al., 2025c ) , MME-RealWorld ( Zhang et al., 2025d )
Hallucination & visual illusions
POPE ( Li et al., 2023 ) , HallusionBench ( Guan et al., 2024 ) , IllusionVQA ( Shahgir et al., 2024 )
Fine-grained perception & chart reading
ChartQA ( Masry et al., 2022 ) , SalBench ( Huynh et al., 2025 )
Table 1: The 25 benchmarks in our evaluation suite, grouped by the capability they primarily target.
No attention intervention
Prior work
Ours
Benchmark
Qwen3-VL 8B-Instruct
Coldstart
No-Sal
VGA
Saliency R1
EASE
Self-Saliency
AlgoPuzzleVQA ↑
32.8
36.9
39.9
34.3
38.7
36.4
41.4
ChartQA ↑
85.0
88.8
89.4
84.9
89.7
88.9
90.0
CV-Bench ↑
86.0
83.4
84.9
86.1
83.8
84.2
85.9
DailyClue ↑
35.4
32.4
34.2
34.2
32.0
32.7
31.2
HallusionBench ↑
63.2
59.7
57.1
62.0
58.3
62.6
55.7
Table 2: Results across 25 benchmarks. Best per benchmark in bold . Mean score is computed over all 25 benchmarks after rescaling MME to [0,100] (raw score /2800 ). Mean rank is the average over per-benchmark ranks (1 = best).
Figure 3: Qualitative examples of how Self-Saliency correctly identifies low-level visual information.
Text-related
Attention-related
Completion
Box union
In-region share
In-region share
length
area
(selected heads)
(layer-22 heads)
Coldstart
230.0
0.533
0.462
0.612
Self-Saliency
149.6
0.566
0.462
0.621
Table 3: Text and attention statistics before and after training, averaged over 100 held-out validation samples. In-region share: attention in grounded regions divided by all visual attention. Self-Saliency primarily adapts its generated text to the selected heads’ attention patterns, while also increasing attention to grounded regions across layer 22.
Model
En border ring
En center
Largest En
Qwen3-VL-8B-Instruct
1.20
0.92
Top left, 6.64
InternVL3.5-8B
1.18
0.95
Bottom left, 2.29
GLM-4.1V-Thinking
1.50
0.82
Top right, 7.79
Nemotron-Omni
1.43
0.87
Top right, 3.86
Human boxes
0.53
1.14
Non-border, 1.87
Table 4: Attention enrichment for different models.
Figure 4: Saliency maps for different models and human-annotated bounding boxes. For each model, each cell represents the fraction of the model’s total visual attention assigned to the corresponding image patch. For the human annotations (right), each cell represents the fraction of images in which the corresponding patch falls within a human-annotated bounding box.
Figure 5: The fraction of the model’s total visual attention assigned to the corresponding image patch, for the heads used to extract saliency (layer 22, heads 28+31), before and after training.
Qwen3-VL-8B Instruct
coldstart
no_sal
center_rect
question_boxes
Self-Saliency
Mean score ( ↑ )
61.14
62.69
63.47
63.39
63.33
64.26
Mean rank ( ↓ )
3.90
4.06
3.34
3.68
3.38
2.64
Wins no. ( ↑ )
7
0
3
2
5
9
Table 5: Performance across our 25-benchmark suite for the ablation study.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
No attention intervention
Prior work
Ours
Benchmark
Qwen3-VL 8B-Instruct
Coldstart
No-Sal
VGA
Saliency- R1
EASE
Self-Saliency
AlgoPuzzleVQA ↑
32.8 ± 1.11
36.9 ± 1.14
39.9 ± 1.15
34.3 ± 1.12
38.7 ± 1.15
36.4 ± 1.13
41.4 ± 1.16
ChartQA ↑
85.0 ± 0.71
88.8 ± 0.63
89.4 ± 0.62
84.9 ± 0.72
89.7 ± 0.61
88.9 ± 0.63
90.0 ± 0.60
CV-Bench ↑
86.0 ± 0.68
83.4 ± 0.72
84.9 ± 0.70
86.1 ± 0.67
83.8 ± 0.72
84.2 ± 0.71
85.9 ± 0.68
DailyClue ↑
35.4 ± 1.85
32.4 ± 1.82
34.2 ± 1.84
34.2 ± 1.84
32.0 ± 1.81
32.7 ± 1.82
31.2 ± 1.80
HallusionBench ↑
63.2 ± 1.56
59.7 ± 1.59
57.1 ± 1.61
62.0 ± 1.57
58.3 ± 1.60
62.6 ± 1.57
55.7 ± 1.57
Appendix
Table 6: Results across 25 benchmarks (mean ± standard error). Best per benchmark in bold . Mean score is computed over all 25 benchmarks after rescaling MME to [0,100] (raw score /2800 ). Mean rank is the average over per-benchmark ranks (1 = best).
No attention intervention
Saliency-based rewards
Benchmark
Split
Qwen3-VL 8B-Instruct
Coldstart
No-Sal
Self-Saliency mean
Self-Saliency
AlgoPuzzleVQA ↑
–
32.8
36.9
39.9
38.8
41.4
ChartQA ↑
–
85.0
88.8
89.4
89.7
90.0
CV-Bench ↑
–
86.0
83.4
84.9
84.4
85.9
DailyClue ↑
–
35.4
32.4
34.2
32.7
31.2
HallusionBench ↑
image
63.2
59.7
57.1
55.7
55.7
Appendix
Table 7: Results across 25 benchmarks, including Self-Saliency mean . Best per benchmark in bold .
Model
Vision encoder
Grid shape
Language model
Qwen3-VL-8B-Instruct
ViT
Image-dependent
Qwen3-8B
InternVL3.5-8B
InternViT-300M
16×16
Qwen3-8B
GLM-4.1V-Thinking
AIMv2-Huge
Image-dependent
GLM-4-9B-0414
Nemotron-Omni
C-RADIOv4-H
Image-dependent
Nemotron
Appendix
Table 8: Models used in the attention focus experiment.
University of Science and Technology of China School of Artificial Intelligence and Data Science Hefei, Anhui, China · University of Science and Technology of China State Key Laboratory of Precision and Intelligent Chemistry Hefei, Anhui, China