We release MoGround, a vision-language dataset spanning four visual domains in which the answer to every question is guaranteed to be available from exactly one modality. This guarantee enables us to measure modality distraction, the failure in which a model answers a question correctly from one modality alone and then flips to a wrong answer once irrelevant content from the other modality is added. Existing probes rarely establish single-modality answerability this way, making it hard to isolate distraction in the first place. Across seven open-source VLMs, we find that modality distraction is not universal but model-dependent. The weaker-grounded modality is the more distracted one (r = +0.86), and distraction scales inversely with grounding strength (r = -0.90). The single-modality guarantee also enables a mitigation method that needs to distinguish between relevant and irrelevant context. Trained on one split of MoGround alone, a weight-space robustness vector reduces distraction on all seven models by 9% to 51%, at a cost of only 0.1 average points of accuracy on standard multimodal tasks.
Figures & tables
Figure 1: The same frozen VLM answers a question twice. Given the image and the question alone, the model answers correctly ( A, green ). Once an irrelevant caption is added, it switches to a wrong answer ( C, red ), distracting the model from the correct option. The answer to the question can be found in the image without any explicit cue in the text, making the query vision-grounded.
Figure 2: A vision-grounded MoGround item from SemArt , answerable via image but not via caption. The color of the flag is visible in the image but nowhere described in the caption.
Accuracy
Distraction Rate
V-distraction rate by domain
Model
(V-grnd)
(T-grnd)
(V-grnd)
(T-grnd)
Photos
Charts
Art
Rad.
Non:Nat Ratio
InternVL3-8B
88
100
3.4
0.3
2.2
4.5
3.8
6.2
2.22×
Qwen2.5-VL-7B
87
100
5.4
0.0
4.6
3.6
5.3
12.7
1.55×
Qwen2.5-VL-3B
82
100
6.4
0.0
4.4
9.6
7.4
8.8
1.95×
LLaVA-OV-7B
78
100
5.9
0.1
3.8
14.0
1.7
11.3
2.40×
Qwen2-VL-2B
71
100
9.1
0.4
10.9
8.2
4.5
8.4
0.65×
Table 1: Distraction results on MoGround-Base. Accuracy is computed with only one modality as input together with the question. V-distraction rate is the fraction of vision-grounded items answered correctly from the image alone but wrongly once a caption is added. Domain breakdown for T-grounded is in Table 8 ). Non:Nat Ratio is all non-photos domains vs photos v-distraction ratio.
Model
V Ground
T Ground
Gap (T − V)
V-Distr
T-Distr
Asymmetry (V − T)
95% CI
InternVL3-8B
96
96
+0.2
1.5
2.3
−0.8 ( ≈ )
[−1.6,+0.0]
LLaVA-OV-7B
99
85
−13.8
0.6
2.3
−1.7 (T > V)
[−2.3,−1.0]
Qwen2.5-VL-7B
90
91
+1.1
2.8
0.7
+2.2 (V > T)
[+1.4,+3.0]
Qwen2.5-VL-3B
90
90
+0.7
2.5
1.5
+1.0 (V > T)
[+0.1,+1.8]
LLaVA-NeXT-8B
96
82
−14.6
2.4
3.6
−1.2 (T > V)
[−2.2,−0.1]
Qwen2-VL-2B
86
74
−11.8
3.0
5.6
−2.6 (T > V)
[−3.9,−1.3]
Table 2: Grounding, distraction, and their asymmetry on the audited MoGround-Retrieved (A-OKVQA Vision / RACE-High Text). Gap = T Ground − V Ground; Asymmetry = V-Distr − T-Distr. The CI is a per-model bootstrap over items. All values are percentages, and Gap and Asymmetry are differences in percentage points. Per-configuration counts and Wilson CIs: Table 9 .
MoGround-Base (Test)
MoGround-Human
MoGround-Retrieved (Held-Out)
Capability (Held-Out)
Backbone
Base
Reduction
Base
Reduction
Base
Reduction
Base
Δ (pp)
Qwen2.5-VL-7B
5.7
−65%
11.3
−4%
2.6
−16%
78
+0.2
InternVL3-8B
2.2
−33%
11.3
−47%
1.6
−38%
83
−0.5
Qwen2.5-VL-3B
5.6
−53%
29.6
−58%
2.8
−20%
72
−0.2
LLaVA-OV-7B
7.3
−73%
13.3
−67%
0.4
−25%
81
+0.2
LLaVA-NeXT-8B
11.5
−47%
40.5
−47%
2.4
−39%
70
−0.5
Table 3: The robustness vector at the default w=0.5 . Each value is the mean over four training seeds. Base is the v-distraction rate at w=0 and Reduction is relative to that base. Capability Δ is an absolute change in accuracy points. The average comes with a 95% interval from 4,000 bootstrap draws that resample items within each backbone
MoGround-Base Test
MoGround-Human
MoGround-Retrieved (Held-Out)
Capability
Method
V-Distr
Rel.
V-Distr
Rel.
V-Distr
Rel.
Δ (pp)
Baseline (No Intervention)
0.076
—
0.274
—
0.026
—
—
Prompt Instruction
0.083
+10%
0.299
+4%
0.029
−1%
+0.3
Zero-Shot Chain-of-Thought
0.120
+57%
0.318
+20%
0.047
+158%
−1.5
M3ID Contrastive Decoding
0.053
−27%
0.229
−19%
0.026
+1%
+0.3
Robustness Vector ( w=0.5 )
0.037
−51%
0.136
−47%
0.019
−29%
−0.1
Table 4: Mean v-distraction and the delta relative to no intervention at all, averaged across four seeds. Capability is the change in general capability. Robustness vector is applied with w=0.5 .
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
MoGround-Base
MoGround-Human
MoGround-Retrieved
Source
Domain
Candidates
Kept
V-grd
T-grd
Total
V-grd
T-grd
Total
V-grd
T-grd
Total
DCI
Natural Photos
5,945
1,736
1052
561
1613
27
26
53
—
—
—
VisText
Statistical Charts
4,782
676
488
188
676
12
12
24
—
—
—
SemArt
Fine-Art Paintings
2,302
546
281
216
497
12
12
24
—
—
—
ROCO
Medical Radiology
1,344
713
255
377
632
12
12
24
—
—
—
A-OKVQA
Natural Photos
—
—
—
—
—
—
—
—
2,261
0
2,261
Appendix
Table 5: Composition of our datasets: MoGround-Base (oracle-verified), MoGround-Human (hand-authored), and MoGround-Retrieved (established VQA and reading-comprehension benchmarks). For MoGround-Base the Candidates and Kept columns give the candidates generated per source and the survivors of the three-condition filter. Kept counts precede deduplication and id renumbering, so they sit slightly above the released V-grd/T-grd totals. MoGround-Retrieved counts are post-audit (Appendix D.2 ).
Annotator 1
Annotator 2
MoGround-Base sample ( 250 items each)
guarantee holds
244 (97.6%)
242 (96.8%)
cross-modal leak
2
4
intended modality insufficient
4
4
source-item defect (mis-key / conflict)
10 / 0
9 / 1
MoGround-Retrieved pool, kept by the filter ( 200 items each; ambiguous-option items excluded)
Appendix
Table 6: The human audit, performed independently by two annotators on the same items, in the same order, under the same protocol. The two agree on 236/250 MoGround-Base items and 189/200 kept MoGround-Retrieved items. The annotators’ instruments split the non-leak categories differently (annotator 1 separated wrong-option support from out-of-set answers. Annotator 2 folded both into conflict), so cross-annotator comparison is at the leak / conflict / defect level. Leak counts are directly comparable.
MoGround-Base
A2: leak
A2: no leak
A1: leak
1
1
A1: no leak
3
245
Appendix
Table 7: Leak / no-leak agreement between the two annotators: MoGround-Base sample (left, n=250 ) and the kept half of MoGround-Retrieved (right, n=200 ; dropped items excluded). Agreement is 98.4% and 99.5% .
T + V
T-only
T-Distraction By Domain
T-Distr
Model
(T-grd)
(T-grd)
Photos
Charts
Art
Rad.
(All)
Qwen2.5-VL-7B
100
100
0.0
0.0
0.0
0.0
0.0
InternVL3-8B
100
100
0.5
0.5
0.0
0.0
0.3
Qwen2.5-VL-3B
100
100
0.0
0.0
0.0
0.0
0.0
LLaVA-OV-7B
100
100
0.2
0.0
0.5
0.0
0.1
LLaVA-NeXT-8B
100
100
0.0
0.0
0.0
0.0
0.0
Appendix
Table 8: Text-grounded half of Table 1 . All values are percentages. T + V and T-only are accuracy on text-grounded items with both inputs and with the caption alone. T-Distraction is the fraction of the T-only-correct items that the added image flips, per domain and pooled. The rate is at or near zero in every configuration, and they are well powered ( n=182 – 557 solved items each), so the ceiling is a property of the verified text side, not of any one domain.
Model
V-Flips
95% CI
T-Flips
95% CI
MoGround-Retrieved (vision and text)
InternVL3-8B
33/2177
[0.011,0.021]
56/2408
[0.018,0.030]
Qwen2.5-VL-7B
58/2036
[0.022,0.037]
15/2274
[0.004,0.011]
Qwen2.5-VL-3B
51/2025
[0.019,0.033]
34/2254
[0.011,0.021]
Qwen2-VL-2B
58/1934
[0.023,0.039]
104/1842
[0.047,0.068]
LLaVA-OV-7B
14/2242
[0.004,0.010]
49/2131
[0.017,0.030]
Appendix
Table 9: Counts and Wilson 95% intervals behind each distraction rate. Each entry reads flipped/solvable : how many items the model answers correctly from the grounded modality alone (solvable), and how many of those it then answers wrongly once the other modality is added (flipped). Top block: MoGround-Retrieved. Bottom block: MoGround-Base with domains pooled.
V-Only
V + T
Net
Items
V-Distr
Recovery
Model
(V-grnd)
(V-grnd)
(pp)
Broken
Fixed
(%)
(%)
InternVL3-8B
88
86
−1.2
61
37
3.4
14.4
Qwen2.5-VL-7B
87
84
−2.8
96
39
5.4
14.1
Qwen2.5-VL-3B
82
79
−3.1
109
45
6.4
12.2
LLaVA-OV-7B
78
76
−1.8
95
58
5.9
12.5
Qwen2-VL-2B
71
68
−3.1
133
69
9.1
11.5
Appendix
Table 10: Both directions of caption-induced answer change on the n=2067 vision-grounded items of MoGround-Base. Broken counts items the model answers from the image alone but misses once the caption is added, and Fixed counts the reverse. V-Distraction is Broken over the image-answerable items, Recovery is Fixed over the rest. All values are percentages except the two counts.
Backbone
n
True Caption
Mismatched
Scrambled
Neutral
Qwen2.5-VL-7B
1790
0.054
0.036
0.039
0.030
InternVL3-8B
1810
0.034
0.028
0.024
0.022
Qwen2.5-VL-3B
1699
0.064
0.037
0.042
0.024
LLaVA-OV-7B
1603
0.059
0.047
0.037
0.021
LLaVA-NeXT-8B
1405
0.108
0.067
0.075
0.048
Qwen2-VL-2B
1466
0.091
0.082
0.071
0.040
Appendix
Table 11: Control captions. Conditional flip rate on each backbone’s V-only-correct items ( n ) when the added text is the item’s true caption (True) vs. a length-matched caption of another same-domain item (Mismatched), the item’s own caption with word order shuffled (Scrambled), or content-free filler of the same length (Neutral). Every contrast is paired on the same items. The true caption is the strongest distractor for all seven backbones.
Backbone (Layer)
Top- K
Recovery (Real)
Recovery (Null)
V/Pass Preserved
Qwen2.5-VL-3B ( L28 )
6
0.037
0.047
0.987
20
0.028
0.019
0.993
50
0.047
0.093
0.987
Dense (Rank-1)
0.037
0.047
0.993
hV Patch (Per-Item)
0.935
0.243
1.000
LLaVA-NeXT-8B ( L18 )
6
0.114
0.127
0.980
Appendix
Table 12: Controllability grid on distracted vision items, both SAE backbones. Recovery is the share of distracted items an edit fixes. The Null column applies the same edit at random. Robust Kept is the share of already-correct items left unchanged.
Intervention
Deployable
Δ V-Distr
Δ Text
Qwen2.5-VL-3B base v-distr 0.056 (19/338 items)
Prompt Instruction
✓
+15%
+0.0
Zero-Shot Chain-of-Thought
✓
+150%
−0.4
M3ID Contrastive Decoding
✓
−22%
+0.0
SAE Feature Ablation ( K=20 )
×
+11%
+0.0
Matched-Random Null
×
−5%
+0.0
Appendix
Table 13: Every intervention on the two backbones that have SAE, on the MoGround-Base test split. “Deployable”: whether a system could run the intervention knowing only the image, the context text, and the question. Δ V-Distr is the relative change in v-distraction from base (negative = reduced). Δ Text is the change in text accuracy in percentage points. Both backbones are scored on identical items on the held-out reporting side.
Intervention
Deployable
Δ V-Distr
Δ Text
Qwen2.5-VL-3B base v-distr 0.028 (28/1009 items)
Prompt Instruction
✓
−14%
−1.0
Zero-Shot Chain-of-Thought
✓
+82%
−12.7
M3ID Contrastive Decoding
✓
+11%
−0.1
SAE Feature Ablation ( K=20 )
×
−4%
+0.0
Matched-Random Null
×
0%
+0.0
Appendix
Table 14: The same interventions on the assembled held-out pool, the cross-domain surface. “Deployable”: whether a system could run the intervention knowing only the image, the context text, and the question. Δ V-Distr is the relative change in v-distraction from base (negative = reduced). Δ Text is the change in text accuracy in percentage points. Both backbones are scored on identical items on the held-out reporting side (Appendix E.6 ). MoGround-Base test is Table 13 .
Intervention
Qwen2.5-VL-3B
LLaVA-NeXT-8B
Prompt Instruction
+0.56
−0.38
Zero-Shot Chain-of-Thought
−2.27
−2.81
Robustness Vector ( w=0.5 )
−0.16
−0.52
Appendix
Table 15: General-capability cost of the three interventions that alter the served model, as the four-benchmark mean accuracy change in percentage points (report-side data). The activation steering methods of Tables 13 and 14 are applied only to items already known to be distracted, and are not models a system could serve, so they have no capability cost to report.
Qwen2.5-VL-3B
LLaVA-NeXT-8B
Intervention
Δ V-Distr
Δ t-Distr
Δ V-Distr
Δ t-Distr
Base Rate
0.028 / 0.020
0.024 / 0.030
Prompt Instruction
−14%
+1.0
0%
−0.1
Zero-Shot Chain-of-Thought
+82%
+15.4
+331%
+16.4
M3ID Contrastive Decoding †
+11%
+0.6
−22%
+2.0
Robustness Vector ( w=0.5 )
−20%
+0.3
−39%
+2.4
Appendix
Table 16: Benefit and cost of the four deployable interventions on the MoGround-Retrieved held-out pool, both measured as conditional rates. Vision side: relative change in v-distraction, negative is better. Text side: change in t-distraction in percentage points, positive is worse. Base Rate gives each backbone’s v- and t-distraction before any intervention.
MoGround-Base (Test)
MoGround-Human
MoGround-Retrieved (Held-Out)
Backbone
Base
TV
Δ (pp)
Base
TV
Δ (pp)
Base
TV
Δ (pp)
Qwen2.5-VL-7B
0.90
0.93
+2.4
0.86
0.87
+1.0
0.90
0.90
−0.1
InternVL3-8B
0.93
0.94
+1.6
0.86
0.87
+0.6
0.95
0.94
−0.6
Qwen2.5-VL-3B
0.86
0.89
+2.7
0.79
0.87
+8.0
0.89
0.89
+0.1
LLaVA-OV-7B
0.85
0.91
+6.4
0.84
0.85
+0.6
0.91
0.91
−0.5
LLaVA-NeXT-8B
0.77
0.82
+5.1
0.72
0.77
+5.4
0.88
0.88
−0.4
Appendix
Table 17: Overall V + T accuracy, base vs. the robustness vector at w=0.5 (4-seed mean), on the three reporting pools: what the conditional reductions of Table 3 convert into end-to-end. Gains track how prevalent distraction is: every backbone gains on both MoGround-Base surfaces, while MoGround-Retrieved, whose baseline distraction is 2 – 3% , is accuracy-neutral on average ( +0.6 pp) with per-model changes within seed noise.
backbone
layer
dmodel
dSAE
FVU
alive frac.
Qwen2.5-VL-3B
L13
2048
16384
0.011
0.64
L20
2048
16384
0.016
0.75
L28
2048
16384
0.009
0.61
L31
2048
16384
0.005
0.40
LLaVA-NeXT-8B
L12
4096
32768
0.081
0.43
L18
4096
32768
0.026
0.40
Appendix
Table 18: Released SAE checkpoints, with reconstruction FVU and alive-feature fraction measured on one evaluation batch.