cs.CVOct 8, 2026

SAGE: Sink-Aware Guided Emphasis for Visual Grounding in Vision-Language Decoders

Authors: Jeonghyo Song, YoungJoon Yoo

Organizations: Department of Artificial Intelligence, Chung-Ang University

Abstract

Recent large vision-language models (VLMs) pair a visual encoder with a large language model (LLM) and perform well on diverse image-text tasks, yet their reliability is often limited by decoder attention pathologies that suppress visual evidence and exacerbate hallucinations. In this paper, we revisit visual attention sinks and uncover a structured, layer-dependent behavior: across prompts, early and late decoder layers exhibit prompt-invariant attention collapse onto the same few image regions, which we term PIS (Prompt-Invariant Sinks), whereas mid layers become prompt-conditioned and drive vision-language alignment. This split suggests that treating sinks as a uniform effect is incomplete. Building on this insight, we propose SAGE (Sink-Aware Guided Emphasis), a lightweight intervention that steers decoder attention away from PIS and toward query-dependent regions of interest (ROIs) using token-aligned ROI masks derived from standard vision backbones such as CLIP, ViT, and DINOv3. Evaluated on diverse vision-encoder + decoder-only LLM VLM families, SAGE improves visual grounding, reduces hallucinations, and yields consistent gains across public downstream vision-language benchmarks, including fine-grained visual discrimination settings where localized evidence is crucial, when instantiated with backbone-derived ROI masks.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models

    Apr 1, 2026Jiho Choi, Jaemin Kim, Sanghwan Kim +2Large Vision-Language ModelsVLM Interpretability

  2. See Only When Needed: Context-Aware Attention Intervention for Mitigating Hallucinations in LVLMs

    Jun 29, 2026Yuqing Lei, Wenbo Lyu, Yingjun Du +3Visual AttentionLarge Vision-Language Models

  3. Instruction-Evidence Contrastive Dual-Stream Decoding for Grounded Vision-Language Reasoning

    Apr 28, 2026Yashwant Pravinrao Bangde, Debaditya RoyObject Hallucination in VLMsVision-Language Grounding