cs.LGSep 27, 2026

MoGround: Measuring and Mitigating Modality Distraction in Vision-Language Models

Authors: Luca Zhou, Bo Zhao, Rose Yu, Emanuele Rodolà, Roberto Dessì

Organizations: Sapienza University of Rome · Harvard University · UC San Diego · Paradigma · Not Diamond

Abstract

We release MoGround, a vision-language dataset spanning four visual domains in which the answer to every question is guaranteed to be available from exactly one modality. This guarantee enables us to measure modality distraction, the failure in which a model answers a question correctly from one modality alone and then flips to a wrong answer once irrelevant content from the other modality is added. Existing probes rarely establish single-modality answerability this way, making it hard to isolate distraction in the first place. Across seven open-source VLMs, we find that modality distraction is not universal but model-dependent. The weaker-grounded modality is the more distracted one (r = +0.86), and distraction scales inversely with grounding strength (r = -0.90). The single-modality guarantee also enables a mitigation method that needs to distinguish between relevant and irrelevant context. Trained on one split of MoGround alone, a weight-space robustness vector reduces distraction on all seven models by 9% to 51%, at a cost of only 0.1 average points of accuracy on standard multimodal tasks.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Are Reasoning Vision-Language Models Robust to Semantic Visual Distractions?

    Jun 8, 2026Yizheng Sun, Mochuan Zhan, Yanan Ma +10Recent Vision-Language ModelsMultimodal Reasoning Benchmarks

  2. Seeing and Solving Are Not Enough for Vision-Language Models

    Sep 27, 2026Ziheng Wang, Mingxuan Xie, Yilin Liu +3Multimodal QueryState-Tracking

  3. Understanding the Effects of Distractors on Reasoning Vision-Language Models

    Nov 26, 2025Jiyun Bae, Hyunjong Ok, Sangwoo Mo +1Recent Vision-Language ModelsDistractor