We release MoGround, a vision-language dataset spanning four visual domains in which the answer to every question is guaranteed to be available from exactly one modality. This guarantee enables us to measure modality distraction, the failure in which a model answers a question correctly from one modality alone and then flips to a wrong answer once irrelevant content from the other modality is added. Existing probes rarely establish single-modality answerability this way, making it hard to isolate distraction in the first place. Across seven open-source VLMs, we find that modality distraction is not universal but model-dependent. The weaker-grounded modality is the more distracted one (r = +0.86), and distraction scales inversely with grounding strength (r = -0.90). The single-modality guarantee also enables a mitigation method that needs to distinguish between relevant and irrelevant context. Trained on one split of MoGround alone, a weight-space robustness vector reduces distraction on all seven models by 9% to 51%, at a cost of only 0.1 average points of accuracy on standard multimodal tasks.
Figures & tables
Figure 1: The same frozen VLM answers a question twice. Given the image and the question alone, the model answers correctly ( A, green ). Once an irrelevant caption is added, it switches to a wrong answer ( C, red ), distracting the model from the correct option. The answer to the question can be found in the image without any explicit cue in the text, making the query vision-grounded.
Figure 2: A vision-grounded MoGround item from SemArt , answerable via image but not via caption. The color of the flag is visible in the image but nowhere described in the caption.
Accuracy
Distraction Rate
V-distraction rate by domain
Model
(V-grnd)
(T-grnd)
(V-grnd)
(T-grnd)
Photos
Charts
Art
Rad.
Non:Nat Ratio
InternVL3-8B
88
100
3.4
0.3
2.2
4.5
3.8
6.2
2.22×
Qwen2.5-VL-7B
87
100
5.4
0.0
4.6
3.6
5.3
12.7
1.55×
Qwen2.5-VL-3B
82
100
6.4
0.0
4.4
9.6
7.4
8.8
1.95×
LLaVA-OV-7B
78
100
5.9
0.1
3.8
14.0
1.7
11.3
2.40×
Qwen2-VL-2B
71
100
9.1
0.4
10.9
8.2
4.5
8.4
0.65×
Table 1: Distraction results on MoGround-Base. Accuracy is computed with only one modality as input together with the question. V-distraction rate is the fraction of vision-grounded items answered correctly from the image alone but wrongly once a caption is added. Domain breakdown for T-grounded is in Table 8 ). Non:Nat Ratio is all non-photos domains vs photos v-distraction ratio.
Model
V Ground
T Ground
Gap (T − V)
V-Distr
T-Distr
Asymmetry (V − T)
95% CI
InternVL3-8B
96
96
+0.2
1.5
2.3
−0.8 ( ≈ )
[−1.6,+0.0]
LLaVA-OV-7B
99
85
−13.8
0.6
2.3
−1.7 (T > V)
[−2.3,−1.0]
Qwen2.5-VL-7B
90
91
+1.1
2.8
0.7
+2.2 (V > T)
[+1.4,+3.0]
Qwen2.5-VL-3B
90
90
+0.7
2.5
1.5
+1.0 (V > T)
[+0.1,+1.8]
LLaVA-NeXT-8B
96
82
−14.6
2.4
3.6
−1.2 (T > V)
[−2.2,−0.1]
Qwen2-VL-2B
86
74
−11.8
3.0
5.6
−2.6 (T > V)
[−3.9,−1.3]
Table 2: Grounding, distraction, and their asymmetry on the audited MoGround-Retrieved (A-OKVQA Vision / RACE-High Text). Gap = T Ground − V Ground; Asymmetry = V-Distr − T-Distr. The CI is a per-model bootstrap over items. All values are percentages, and Gap and Asymmetry are differences in percentage points. Per-configuration counts and Wilson CIs: Table 9 .
MoGround-Base (Test)
MoGround-Human
MoGround-Retrieved (Held-Out)
Capability (Held-Out)
Backbone
Base
Reduction
Base
Reduction
Base
Reduction
Base
Δ (pp)
Qwen2.5-VL-7B
5.7
−65%
11.3
−4%
2.6
−16%
78
+0.2
InternVL3-8B
2.2
−33%
11.3
−47%
1.6
−38%
83
−0.5
Qwen2.5-VL-3B
5.6
−53%
29.6
−58%
2.8
−20%
72
−0.2
LLaVA-OV-7B
7.3
−73%
13.3
−67%
0.4
−25%
81
+0.2
LLaVA-NeXT-8B
11.5
−47%
40.5
−47%
2.4
−39%
70
−0.5
Table 3: The robustness vector at the default w=0.5 . Each value is the mean over four training seeds. Base is the v-distraction rate at w=0 and Reduction is relative to that base. Capability Δ is an absolute change in accuracy points. The average comes with a 95% interval from 4,000 bootstrap draws that resample items within each backbone
MoGround-Base Test
MoGround-Human
MoGround-Retrieved (Held-Out)
Capability
Method
V-Distr
Rel.
V-Distr
Rel.
V-Distr
Rel.
Δ (pp)
Baseline (No Intervention)
0.076
—
0.274
—
0.026
—
—
Prompt Instruction
0.083
+10%
0.299
+4%
0.029
−1%
+0.3
Zero-Shot Chain-of-Thought
0.120
+57%
0.318
+20%
0.047
+158%
−1.5
M3ID Contrastive Decoding
0.053
−27%
0.229
−19%
0.026
+1%
+0.3
Robustness Vector ( w=0.5 )
0.037
−51%
0.136
−47%
0.019
−29%
−0.1
Table 4: Mean v-distraction and the delta relative to no intervention at all, averaged across four seeds. Capability is the change in general capability. Robustness vector is applied with w=0.5 .
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
MoGround-Base
MoGround-Human
MoGround-Retrieved
Source
Domain
Candidates
Kept
V-grd
T-grd
Total
V-grd
T-grd
Total
V-grd
T-grd
Total
DCI
Natural Photos
5,945
1,736
1052
561
1613
27
26
53
—
—
—
VisText
Statistical Charts
4,782
676
488
188
676
12
12
24
—
—
—
SemArt
Fine-Art Paintings
2,302
546
281
216
497
12
12
24
—
—
—
ROCO
Medical Radiology
1,344
713
255
377
632
12
12
24
—
—
—
A-OKVQA
Natural Photos
—
—
—
—
—
—
—
—
2,261
0
2,261
Appendix
Table 5: Composition of our datasets: MoGround-Base (oracle-verified), MoGround-Human (hand-authored), and MoGround-Retrieved (established VQA and reading-comprehension benchmarks). For MoGround-Base the Candidates and Kept columns give the candidates generated per source and the survivors of the three-condition filter. Kept counts precede deduplication and id renumbering, so they sit slightly above the released V-grd/T-grd totals. MoGround-Retrieved counts are post-audit (Appendix D.2 ).
Annotator 1
Annotator 2
MoGround-Base sample ( 250 items each)
guarantee holds
244 (97.6%)
242 (96.8%)
cross-modal leak
2
4
intended modality insufficient
4
4
source-item defect (mis-key / conflict)
10 / 0
9 / 1
MoGround-Retrieved pool, kept by the filter ( 200 items each; ambiguous-option items excluded)
Appendix
Table 6: The human audit, performed independently by two annotators on the same items, in the same order, under the same protocol. The two agree on 236/250 MoGround-Base items and 189/200 kept MoGround-Retrieved items. The annotators’ instruments split the non-leak categories differently (annotator 1 separated wrong-option support from out-of-set answers. Annotator 2 folded both into conflict), so cross-annotator comparison is at the leak / conflict / defect level. Leak counts are directly comparable.
MoGround-Base
A2: leak
A2: no leak
A1: leak
1
1
A1: no leak
3
245
Appendix
Table 7: Leak / no-leak agreement between the two annotators: MoGround-Base sample (left, n=250 ) and the kept half of MoGround-Retrieved (right, n=200 ; dropped items excluded). Agreement is 98.4% and 99.5% .
T + V
T-only
T-Distraction By Domain
T-Distr
Model
(T-grd)
(T-grd)
Photos
Charts
Art
Rad.
(All)
Qwen2.5-VL-7B
100
100
0.0
0.0
0.0
0.0
0.0
InternVL3-8B
100
100
0.5
0.5
0.0
0.0
0.3
Qwen2.5-VL-3B
100
100
0.0
0.0
0.0
0.0
0.0
LLaVA-OV-7B
100
100
0.2
0.0
0.5
0.0
0.1
LLaVA-NeXT-8B
100
100
0.0
0.0
0.0
0.0
0.0
Appendix
Table 8: Text-grounded half of Table 1 . All values are percentages. T + V and T-only are accuracy on text-grounded items with both inputs and with the caption alone. T-Distraction is the fraction of the T-only-correct items that the added image flips, per domain and pooled. The rate is at or near zero in every configuration, and they are well powered ( n=182 – 557 solved items each), so the ceiling is a property of the verified text side, not of any one domain.
Model
V-Flips
95% CI
T-Flips
95% CI
MoGround-Retrieved (vision and text)
InternVL3-8B
33/2177
[0.011,0.021]
56/2408
[0.018,0.030]
Qwen2.5-VL-7B
58/2036
[0.022,0.037]
15/2274
[0.004,0.011]
Qwen2.5-VL-3B
51/2025
[0.019,0.033]
34/2254
[0.011,0.021]
Qwen2-VL-2B
58/1934
[0.023,0.039]
104/1842
[0.047,0.068]
LLaVA-OV-7B
14/2242
[0.004,0.010]
49/2131
[0.017,0.030]
Appendix
Table 9: Counts and Wilson 95% intervals behind each distraction rate. Each entry reads flipped/solvable : how many items the model answers correctly from the grounded modality alone (solvable), and how many of those it then answers wrongly once the other modality is added (flipped). Top block: MoGround-Retrieved. Bottom block: MoGround-Base with domains pooled.
V-Only
V + T
Net
Items
V-Distr
Recovery
Model
(V-grnd)
(V-grnd)
(pp)
Broken
Fixed
(%)
(%)
InternVL3-8B
88
86
−1.2
61
37
3.4
14.4
Qwen2.5-VL-7B
87
84
−2.8
96
39
5.4
14.1
Qwen2.5-VL-3B
82
79
−3.1
109
45
6.4
12.2
LLaVA-OV-7B
78
76
−1.8
95
58
5.9
12.5
Qwen2-VL-2B
71
68
−3.1
133
69
9.1
11.5
Appendix
Table 10: Both directions of caption-induced answer change on the n=2067 vision-grounded items of MoGround-Base. Broken counts items the model answers from the image alone but misses once the caption is added, and Fixed counts the reverse. V-Distraction is Broken over the image-answerable items, Recovery is Fixed over the rest. All values are percentages except the two counts.
Backbone
n
True Caption
Mismatched
Scrambled
Neutral
Qwen2.5-VL-7B
1790
0.054
0.036
0.039
0.030
InternVL3-8B
1810
0.034
0.028
0.024
0.022
Qwen2.5-VL-3B
1699
0.064
0.037
0.042
0.024
LLaVA-OV-7B
1603
0.059
0.047
0.037
0.021
LLaVA-NeXT-8B
1405
0.108
0.067
0.075
0.048
Qwen2-VL-2B
1466
0.091
0.082
0.071
0.040
Appendix
Table 11: Control captions. Conditional flip rate on each backbone’s V-only-correct items ( n ) when the added text is the item’s true caption (True) vs. a length-matched caption of another same-domain item (Mismatched), the item’s own caption with word order shuffled (Scrambled), or content-free filler of the same length (Neutral). Every contrast is paired on the same items. The true caption is the strongest distractor for all seven backbones.
Backbone (Layer)
Top- K
Recovery (Real)
Recovery (Null)
V/Pass Preserved
Qwen2.5-VL-3B ( L28 )
6
0.037
0.047
0.987
20
0.028
0.019
0.993
50
0.047
0.093
0.987
Dense (Rank-1)
0.037
0.047
0.993
hV Patch (Per-Item)
0.935
0.243
1.000
LLaVA-NeXT-8B ( L18 )
6
0.114
0.127
0.980
Appendix
Table 12: Controllability grid on distracted vision items, both SAE backbones. Recovery is the share of distracted items an edit fixes. The Null column applies the same edit at random. Robust Kept is the share of already-correct items left unchanged.
Intervention
Deployable
Δ V-Distr
Δ Text
Qwen2.5-VL-3B base v-distr 0.056 (19/338 items)
Prompt Instruction
✓
+15%
+0.0
Zero-Shot Chain-of-Thought
✓
+150%
−0.4
M3ID Contrastive Decoding
✓
−22%
+0.0
SAE Feature Ablation ( K=20 )
×
+11%
+0.0
Matched-Random Null
×
−5%
+0.0
Appendix
Table 13: Every intervention on the two backbones that have SAE, on the MoGround-Base test split. “Deployable”: whether a system could run the intervention knowing only the image, the context text, and the question. Δ V-Distr is the relative change in v-distraction from base (negative = reduced). Δ Text is the change in text accuracy in percentage points. Both backbones are scored on identical items on the held-out reporting side.
Intervention
Deployable
Δ V-Distr
Δ Text
Qwen2.5-VL-3B base v-distr 0.028 (28/1009 items)
Prompt Instruction
✓
−14%
−1.0
Zero-Shot Chain-of-Thought
✓
+82%
−12.7
M3ID Contrastive Decoding
✓
+11%
−0.1
SAE Feature Ablation ( K=20 )
×
−4%
+0.0
Matched-Random Null
×
0%
+0.0
Appendix
Table 14: The same interventions on the assembled held-out pool, the cross-domain surface. “Deployable”: whether a system could run the intervention knowing only the image, the context text, and the question. Δ V-Distr is the relative change in v-distraction from base (negative = reduced). Δ Text is the change in text accuracy in percentage points. Both backbones are scored on identical items on the held-out reporting side (Appendix E.6 ). MoGround-Base test is Table 13 .
Intervention
Qwen2.5-VL-3B
LLaVA-NeXT-8B
Prompt Instruction
+0.56
−0.38
Zero-Shot Chain-of-Thought
−2.27
−2.81
Robustness Vector ( w=0.5 )
−0.16
−0.52
Appendix
Table 15: General-capability cost of the three interventions that alter the served model, as the four-benchmark mean accuracy change in percentage points (report-side data). The activation steering methods of Tables 13 and 14 are applied only to items already known to be distracted, and are not models a system could serve, so they have no capability cost to report.
Qwen2.5-VL-3B
LLaVA-NeXT-8B
Intervention
Δ V-Distr
Δ t-Distr
Δ V-Distr
Δ t-Distr
Base Rate
0.028 / 0.020
0.024 / 0.030
Prompt Instruction
−14%
+1.0
0%
−0.1
Zero-Shot Chain-of-Thought
+82%
+15.4
+331%
+16.4
M3ID Contrastive Decoding †
+11%
+0.6
−22%
+2.0
Robustness Vector ( w=0.5 )
−20%
+0.3
−39%
+2.4
Appendix
Table 16: Benefit and cost of the four deployable interventions on the MoGround-Retrieved held-out pool, both measured as conditional rates. Vision side: relative change in v-distraction, negative is better. Text side: change in t-distraction in percentage points, positive is worse. Base Rate gives each backbone’s v- and t-distraction before any intervention.
MoGround-Base (Test)
MoGround-Human
MoGround-Retrieved (Held-Out)
Backbone
Base
TV
Δ (pp)
Base
TV
Δ (pp)
Base
TV
Δ (pp)
Qwen2.5-VL-7B
0.90
0.93
+2.4
0.86
0.87
+1.0
0.90
0.90
−0.1
InternVL3-8B
0.93
0.94
+1.6
0.86
0.87
+0.6
0.95
0.94
−0.6
Qwen2.5-VL-3B
0.86
0.89
+2.7
0.79
0.87
+8.0
0.89
0.89
+0.1
LLaVA-OV-7B
0.85
0.91
+6.4
0.84
0.85
+0.6
0.91
0.91
−0.5
LLaVA-NeXT-8B
0.77
0.82
+5.1
0.72
0.77
+5.4
0.88
0.88
−0.4
Appendix
Table 17: Overall V + T accuracy, base vs. the robustness vector at w=0.5 (4-seed mean), on the three reporting pools: what the conditional reductions of Table 3 convert into end-to-end. Gains track how prevalent distraction is: every backbone gains on both MoGround-Base surfaces, while MoGround-Retrieved, whose baseline distraction is 2 – 3% , is accuracy-neutral on average ( +0.6 pp) with per-model changes within seed noise.
backbone
layer
dmodel
dSAE
FVU
alive frac.
Qwen2.5-VL-3B
L13
2048
16384
0.011
0.64
L20
2048
16384
0.016
0.75
L28
2048
16384
0.009
0.61
L31
2048
16384
0.005
0.40
LLaVA-NeXT-8B
L12
4096
32768
0.081
0.43
L18
4096
32768
0.026
0.40
Appendix
Table 18: Released SAE checkpoints, with reconstruction FVU and alive-feature fraction measured on one evaluation batch.
Reasoning Vision-Language Models (VLMs) achieve strong performance on complex multimodal tasks, but reliable real-world application requires handling visual inputs that are messier than clean, curated benchmarks. Existing works mainly evaluate such reliability of VLMs through input corruptions, such as noise, blur and weather effects, which make visual evidence harder to perceive. This leaves a critical reliability failure mode underexplored: a model may perceive the evidence correctly, yet reason from plausible but irrelevant and distracting evidence and propagate this mistake to its final answer. To address this gap, we introduce \textbf{Distract-Bench}, a benchmark for evaluating VLM robustness to \textbf{semantic visual distractions}, defined as meaningful but task-irrelevant visual cues added to inputs while preserving the ground-truth answer. We comprehensively evaluate eight leading open-source and two closed-source VLMs across conventional vision corruptions and Distract-Bench. Our results show that Distract-Bench exposes a robustness failure distinct from vision corruptions: reasoning VLMs largely track their non-reasoning base models under perceptual degradation, but show consistently lower robustness to semantic distractions. Further analysis shows that these distractions often enter the reasoning process of VLMs, are treated as evidence, and lead to incorrect answers. Together, these findings reframe robustness evaluation for reasoning VLMs, shifting the focus from degraded perception to distractions for reliable real-world visual reasoning. Our data and code are available at https://github.com/Yizheng-Sun/Distract-Bench.
Yizheng Sun, Mochuan Zhan, Yanan Ma +10
University of Manchester · Marex · Imperial College London
Vision-language models (VLMs) answer visual questions by combining visual information extraction with downstream problem solving. We investigate a fundamental question: Does an incorrect answer necessarily reflect a failure in visual extraction or problem solving? A model may succeed at both abilities when tested separately yet still fail on the original multimodal question, a distinction that overall answer accuracy cannot reveal. To study this, we perform a question-level empirical analysis across multiple VLMs and visual domains. We define an exactly scorable task state (i.e., the visual information sufficient to solve a question) and use it to test whether the same model can extract the required state, solve the question from the ground-truth state, and answer the original multimodal question. We find that composition failures, where extraction and solving both succeed but direct answering fails, account for 17.7% to 75.6% of direct-answering errors across multiple VLMs and datasets. To address this failure mode, we introduce a simple yet effective method, termed State Realization Tuning (SRT). SRT fine-tunes LoRA adapters attached to the language-model layers while keeping the pretrained VLM weights frozen. It trains the model to output the ground-truth task state before the final answer in a single autoregressive response. SRT improves over standard supervised fine-tuning by 1.7 to 14.1 percentage points and repairs 92.5% to 98.1% of diagnosed composition failures. A single LoRA adapter trained with SRT also improves performance across substantially different task-state structures. Our work shows that having both visual extraction and problem-solving capabilities does not guarantee correct multimodal answering. Requiring the model to first output the visual information needed to solve the question can help bridge this gap.
Ziheng Wang, Mingxuan Xie, Yilin Liu +3
Sun Yat-sen University · Zhejiang University · The Hong Kong University of Science and Technology +2
How does irrelevant information (i.e., distractors) affect test-time scaling in vision-language models (VLMs)? Prior work on text-only language models has shown that textual distractors can intensify inverse scaling, causing models to reason longer but less effective reasoning traces. In this work, we investigate whether similar phenomena arise in multimodal settings. We introduce Idis (Images with distractors), a visual question-answering dataset that systematically varies distractors along semantic and numerical dimensions. Our analyses reveal that visual distractors affect reasoning VLMs in a fundamentally different way from textual distractors: although inverse scaling still emerges, visual distractors reduce accuracy without increasing reasoning length. We further show that attribute counts extracted from reasoning traces provide key insights into how distractors interact with reasoning length and accuracy. As a sanity check, we propose a simple prompting strategy that mitigates distractor-driven predictions in reasoning vision-language models.
Jiyun Bae, Hyunjong Ok, Sangwoo Mo +1
Pohang University of Science and Technology (POSTECH)