SAGE: Sink-Aware Guided Emphasis for Visual Grounding in Vision-Language Decoders
Organizations: Department of Artificial Intelligence, Chung-Ang University
Abstract
Recent large vision-language models (VLMs) pair a visual encoder with a large language model (LLM) and perform well on diverse image-text tasks, yet their reliability is often limited by decoder attention pathologies that suppress visual evidence and exacerbate hallucinations. In this paper, we revisit visual attention sinks and uncover a structured, layer-dependent behavior: across prompts, early and late decoder layers exhibit prompt-invariant attention collapse onto the same few image regions, which we term PIS (Prompt-Invariant Sinks), whereas mid layers become prompt-conditioned and drive vision-language alignment. This split suggests that treating sinks as a uniform effect is incomplete. Building on this insight, we propose SAGE (Sink-Aware Guided Emphasis), a lightweight intervention that steers decoder attention away from PIS and toward query-dependent regions of interest (ROIs) using token-aligned ROI masks derived from standard vision backbones such as CLIP, ViT, and DINOv3. Evaluated on diverse vision-encoder + decoder-only LLM VLM families, SAGE improves visual grounding, reduces hallucinations, and yields consistent gains across public downstream vision-language benchmarks, including fine-grained visual discrimination settings where localized evidence is crucial, when instantiated with backbone-derived ROI masks.
Figures & tables
| Model | TextVQA | POPE | GQA | ScienceQA | MME | MMVP | MMVP Original | MMVet | |||
| Acc | F1 | Image | Pair | Image | Pair | ||||||
| LLaVA 1.5-13B | 55.50 | 82.75 | 81.55 | 61.65 | 63.98 | 1783.11 | 61.48 | 11.85 | 33.58 | 10.65 | 31.20 |
| 61.94 | 86.84 | 86.58 | 62.65 | 68.63 | 1842.58 | 62.96 | 37.78 | 50.00 | 22.00 | 42.20 | |
| + SAGE (CLIP) | ( +6.44 ) | ( +4.09 ) | ( +5.03 ) | ( +1.00 ) | ( +4.65 ) | ( +59.47 ) | ( +1.48 ) | ( +25.93 ) | ( +16.42 ) | ( +11.35 ) | ( +11.00 ) |
| 61.85 | 85.81 | 84.82 | 65.50 | 69.15 | 1895.55 | 68.89 | 38.52 | 58.58 | 21.27 | 34.50 | |
| + SAGE (ViT) | ( +6.35 ) | ( +3.06 ) | ( +3.27 ) | ( +3.85 ) | ( +5.17 ) | ( +112.44 ) | ( +7.41 ) | ( +26.67 ) | ( +25.00 ) | ( +10.62 ) | ( +3.30 ) |
| Depth composition | Mean Acc. |
| 2-layer sets | |
| Very-early only | 19.19 |
| Early-mixed | 49.85 |
| Mid–mid | 61.96 |
| Mid–final | 62.60 |
| Final–final | 62.42 |
| SAGE | ||||
| Benchmark | Base | Best | L24 | Rand. |
| TextVQA | 44.0 | 55.4 | 54.6 | 43.1 |
| GQA | 57.1 | 62.8 | 62.3 | 57.0 |
| POPE | 72.7 | 73.6 | 73.1 | 72.1 |
| POPE | 69.3 | 73.0 | 69.9 | 69.1 |
| ScienceQA | 68.1 | 68.6 | 68.6 | 68.0 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Value / Description |
| Local VLM | Vision-encoder + decoder-only LLM VLMs (e.g., LLaVA-1.5 Liu et al. (2024b) /Qwen2-VL Wang et al. (2024) /InternVL3 Zhu et al. (2025) /NVILA Liu et al. (2025) ) with standard preprocessing and decoding settings. |
| ROI sources | Patch-level ROI masks derived from standard pretrained vision backbones (CLIP/ViT/DINOv3). |
| Patch relevance scoring | Backbone-specific patch relevance scores are thresholded to form a binary ROI mask; the same scoring rule is used consistently within each source. |
| Patch-to-token mapping | The patch-level mask is aligned to the visual-token grid used by the evaluated VLM; the mapping is fixed per model and preprocessing configuration. |
| Patch grid | Model-specific patch grid size determined by the vision encoder and preprocessing; grid configuration is fixed per model and recorded in the experiment scripts. |
| Intervention layers | Layer set specified as a list or range (single-layer or multi-layer), fixed per experiment. |
| Variant | MMVP (%) | |
| No re-weighting | 61.48 | |
| Pull-only | 54.07 | |
| Push-only | 57.49 | |
| Weak | 61.91 | |
| Default | 62.96 | |
| Strong | 59.39 |
| Model | TextVQA | POPE | GQA | SQA | MME | MMVet | MMVP Img | MMVP Pair | MMVP-O Img | MMVP-O Pair |
| LLaVA 1.5-13B | ||||||||||
| + SAGE (CLIP) | L17 | L3 | L20 | L18 | L24 | L20 | L26 | L26 | L18 | L6 |
| + SAGE (ViT) | L29 | L10 | L19 | L20 | L22 | L18 | L18 | L18 | L2 | L18 |
| + SAGE (DINOv3) | L30 | L10 | L30 | L25 | L16 | L20 | L26 | L26 | L18 | L18 |
| Qwen2-VL-7B | ||||||||||
| + SAGE (CLIP) | L11 | L18 | L21 | L15 | L25 | L14 | L18 | L18 | L15 | L15 |
| Layer set | Depth composition | Accuracy (%) |
| (14,16) | Mid–mid | 62.43 |
| (17,21) | Mid–final | 62.72 |
| (28,31) | Final–final | 62.58 |
| Layer set | Depth composition | Accuracy (%) |
| (14,18,21) | Mid | 62.33 |
| (17,24,30) | Mid–final–final | 62.63 |
| (24,27,30) | Final | 62.53 |
| Layer set | Depth composition | Accuracy (%) |
| (0,2) | Very-early only | 12.07 |
| (0,2,4) | Very-early | 7.12 |
| (1,2,3) | Very-early | 1.12 |