As vision language models are increasingly deployed in clinical diagnosis, understanding how they internally resolve competing visual and textual signals becomes a safety imperative. Existing mechanistic analyses remain confined to unimodal text and offer no explanation for why a single misleading sentence can override a correct image based diagnosis, or why a model commits to a confident answer despite insufficient visual evidence. We find that these two safety risks, arbitration failure where textual context overrides visual grounding and brake failure where the model commits without adequate evidence, are mediated by spatially disjoint attention head populations: arbitration heads form a mid-to-deep wideband reflecting cross-layer evidence competition, while brake heads concentrate in a narrow middle-to-late layer band that regulates evidence sufficiency and abstention behavior. To ground these observations in causal circuitry, we introduce CRAFT, which localizes each failure mode to a minimal causal head set via dual criteria and verifies necessity and sufficiency through temporal probes and Tuned Lens trajectory analysis. Excising arbitration heads sharply reduces conflict following with negligible degradation on clean inputs, while excising brake heads restores appropriate abstention under degraded visual evidence. The two interventions target spatially disjoint head sets and produce distinct corrective effects, underscoring the mechanistic separability of the failure modes. Experiments across multiple medical VQA benchmarks and VLM architectures validate both the localization and interventions, demonstrating that the identified heads causally drive each failure mode and that targeted modulation generalises without retraining. The code is available at GitHub repository.
Figures & tables
Figure 1 : Two failure modes in medical VLMs and Craft interventions. Left (Arbitration Failure): Contradictory text hijacks a correct visual diagnosis, causing the model to follow the misleading answer. Right (Brake Failure): The model commits despite occluded evidence.
Figure 2 : Interventional tracing setup. (a) The evidence conflict problem can be modeled by a structural causal model with three competing paths: visual anchoring Pv , parametric knowledge PΘ , and textual override Pt , plus latent confounders Uc that open back-door paths. (b) Path dominance under each failure mode. (c) Clean-state interchange blocks back-door paths through Uc while preserving the front-door visual causal path, isolating genuine causal effects for head selection.
Figure 3 : Arbitration failure localization. (a) Layer-wise S(l) : shaded window marks layers where override concentrates. (b) CER vs. BCP: red points (upper left) are selected causal key heads.
Figure 4 : Arbitration head localization (Hulu-Med 4B). Left: conflict heads (blue) cluster in the Override Band with high CER and low BCP; the Late Commitment Tail shows negligible conflict or fidelity circuitry, with residual score reflecting distributed output formation. Centre: layer × head grid with per layer CER and BCP bars. Right: selection criterion and recovery example.
Figure 5 : Brake head localization (Hulu-Med 4B). (a) HR( l ) map with positive peaks at the deepest layers. (b,c) BCP versus HR landscape separating different heads from backbone.
Prefix
Before Q
Before A
Benchmark
Model
Method
CFR ↓
Resist ↑
C2W ↓
W2C ↑
CFR ↓
Resist ↑
C2W ↓
W2C ↑
CFR ↓
Resist ↑
C2W ↓
W2C ↑
VQA-RAD
Hulu-Med 4B
Baseline
33.33
66.67
N/A
N/A
44.62
55.38
N/A
N/A
98.21
1.79
N/A
N/A
Random
33.59
66.41
1.79
1.54
40.77
59.23
1.03
4.87
98.72
1.28
0.51
0.00
Craft
15.38
84.62
0.26
18.21
14.36
85.64
0.51
30.77
70.26
29.74
0.00
27.95
Hulu-Med 7B
Baseline
25.80
74.20
N/A
N/A
36.50
63.50
N/A
N/A
88.00
12.00
N/A
N/A
Random
25.40
74.60
0.90
1.30
35.20
64.80
0.80
2.10
88.30
11.70
0.35
0.05
Table 1 : Arbitration failure on multimodal benchmarks under textual conflict. CFR : conflict following rate ( ↓ ); Resist: resist rate ( ↑ ); C2W : correct to wrong ( ↓ ); W2C: wrong to correct ( ↑ ). Bold: best comparison result. All CRAFT gains are significant ( p<0.01 ).
Hulu-Med 4B
Hulu-Med 7B
Hulu-Med 14B
InternVL3.5 4B
Qwen3-VL-8B
Benchmark
Method
UR ↑
O2U ↑
U2O ↓
UR ↑
O2U ↑
U2O ↓
UR ↑
O2U ↑
U2O ↓
UR ↑
O2U ↑
U2O ↓
UR ↑
O2U ↑
U2O ↓
SLAKE
Baseline
37.88
–
–
44.12
–
–
48.90
–
–
22.58
–
–
42.42
–
–
Random
45.45
8.33
0.76
48.82
5.29
0.59
51.10
3.00
0.80
16.13
0.00
6.45
47.73
6.06
0.75
Craft
65.15
27.27
0.00
61.18
17.06
0.00
58.40
10.20
0.70
32.26
9.68
0.00
55.30
12.88
0.00
HealMed-VQA
Baseline
78.44
–
–
85.07
–
–
88.61
–
–
81.56
–
–
84.72
–
–
Random
86.83
10.18
1.80
88.19
4.31
1.19
89.58
1.45
0.48
82.47
3.06
2.15
87.50
4.17
1.39
Table 2 : Brake failure under visual degradation. Craft excises the localized brake heads to recover abstention. UR: unknown rate ( ↑ ); O2U: non-unknown to unknown ( ↑ ); U2O: unknown to other ( ↓ ).
Model
Criterion
CFR↓
∣ΔCFR∣↑
C2W↓
#Heads
Hulu-Med 4B
CER only
6.7
37.9
1.8
6
BCP only
44.4
0.3
1.3
6
Dual (Ours)
14.4
30.3
0.5
6
InternVL3.5 4B
CER only
9.7
12.0
4.0
10
BCP only
7.4
14.3
2.3
10
Dual (Ours)
7.4
14.3
0.0
10
Table 3 : Comparison of selection criteria at equal head count. Left: arbitration failure (CER only: highest conflict effect; BCP only: lowest collateral). Right: brake failure on validation split (HR only: highest hallucination relief; BCP only: lowest collateral). Shaded: CRAFT.
Figure 6 : Temporal probe separation. Probe A (detection) saturates before Probe B (commitment), revealing a detect then commit structure. Post-intervention selectively suppresses Probe B while preserving Probe A.
Figure 7 : Tuned Lens trajectories. The preference margin Δl flips sign within the override window under conflict. Intervention suppresses this reversal, restoring correct dominance.
Figure 8 : The Unknown rate increases with mask_scale , but remains below 50% even under the strongest masking. We use mask_scale =2.0 as the default degradation setting.
τs
τd
CFR ↓
C2W ↓
W2C ↑
#Heads
0.133
0.550
18.7
4.9
30.8
25
0.271
0.183
24.6
0.5
20.5
2
0.156
0.445
22.1
7.9
30.5
31
0.350
0.500
18.2
0.8
27.4
6
Table 4 : Threshold sensitivity (Hulu-Med 4B). Left : arbitration on VQA-RAD (Before Q). Right : brake on SLAKE (image conflict). Highlighted: final settings.
ConflictMedQA
PubMedQA
Position
Model
Method
CFR ↓
Resist ↑
C2W ↓
W2C ↑
CFR ↓
Resist ↑
C2W ↓
W2C ↑
Prefix
Qwen3-4B
Baseline
83.33
16.67
–
–
49.67
50.33
–
–
Random
94.25
5.75
12.07
1.15
57.17
42.83
12.17
4.67
Craft
37.36
62.64
0.00
45.98
23.59
76.41
0.00
26.09
Llama3.2-3B
Baseline
57.69
42.31
–
–
58.46
41.54
–
–
Random
42.31
57.69
26.92
42.31
63.08
36.92
8.85
4.23
Table 5 : Text only benchmark results. CFR : conflict compliance ( ↓ ); Resist : resist rate ( ↑ ); C2W : correct to wrong ( ↓ ); W2C : wrong to correct ( ↑ ). Bold : best per block; shaded rows: Craft . Random uses the same head count.
Method
Paradigm
Extra Model
Training
Target Failure
CCD [ 41 ]
Decoding guidance
✓
✗
Textual conflict
MARINE [ 42 ]
Decoding guidance
✓
✗
Textual conflict
ITI [ 20 ]
Activation steering
✗
✓
General truthfulness
HALP [ 17 ]
Hallucination detection
✗
✓
Pre generation risk
Craft
Causal head repair
✗
✗
Arbitration + brake
Table 6 : Inference time safety baselines. Highlighted: the proposed method.
Table 9 : Clean accuracy preservation under baseline methods. Δ Acc denotes post minus pre intervention accuracy (%). Bold: best preservation result.
Method
Training
CFR↓
Resist ↑
C2W ↓
W2C ↑
Clean Acc. ↑
Baseline
✗
44.62
55.38
0.00
0.00
100.00
Random heads
✗
40.77
59.23
1.03
4.87
96.67
Conflict SFT
✓
22.31
77.69
2.05
24.36
97.18
Preference Tuning
✓
19.49
80.51
1.54
26.67
97.44
Craft
✗
14.36
85.64
0.51
30.77
96.41
Conflict SFT + Craft
✓
10.77
89.23
1.03
35.90
96.92
Table 10 : Comparison with fine tuning based safety adaptation on textual conflict. Hulu-Med 4B, VQA-RAD Before Q. Shaded rows denote Craft .
Method
Training
UR↑
O2U ↑
U2O ↓
Clean Acc. ↑
Net Gain ↑
Baseline
✗
37.88
0.00
0.00
100.00
0.00
Random heads
✗
45.45
8.33
0.76
96.90
7.57
Conflict SFT
✓
54.55
19.70
2.27
96.97
16.67
Preference Tuning
✓
57.58
21.97
1.52
96.21
19.70
Craft
✗
65.15
27.27
0.00
90.67
27.27
Conflict SFT + Craft
✓
69.70
32.58
0.76
92.42
31.82
Table 11 : Comparison with fine tuning based safety adaptation on visual brake failure. Hulu-Med 4B, SLAKE image conflict pool. Shaded rows denote Craft .
Hulu-Med 4B
InternVL3.5 4B
Benchmark
Method
Gold ↑
Conflict ↓
Qualified ↑
Clean Fact. ↑
Gold ↑
Conflict ↓
Qualified ↑
Clean Fact. ↑
VQA RAD
Baseline
54.20
39.60
6.20
96.00
56.40
36.80
6.80
96.30
Random
57.10
35.70
7.20
95.10
58.20
34.70
7.10
95.40
Craft
70.90
21.80
7.30
94.70
78.40
13.20
8.40
95.00
SLAKE
Baseline
51.80
42.20
6.00
95.30
59.10
34.60
6.30
96.80
Random
54.40
38.60
7.00
94.60
60.30
32.20
7.50
96.10
Table 12 : Open ended textual conflict results. The judge classifies each response into gold claim, conflict claim, or qualified response. Shaded rows denote Craft .
Hulu-Med 4B
InternVL3.5 4B
Benchmark
Method
Uncert. ↑
Unsup. ↓
Evid. Req. ↑
Clean Compl. ↑
Uncert. ↑
Unsup. ↓
Evid. Req. ↑
Clean Compl. ↑
SLAKE
Baseline
36.20
59.10
4.70
95.60
24.80
69.40
5.80
94.80
Random
44.00
50.70
5.30
94.70
18.60
75.20
6.20
93.70
Craft
63.40
29.90
6.70
90.80
34.80
58.60
6.60
92.80
HealMed-VQA
Baseline
73.60
21.80
4.60
96.20
76.40
18.70
4.90
95.60
Random
79.10
15.80
5.10
95.10
78.80
16.00
5.20
94.70
Table 13 : Open ended degraded vision results. The judge evaluates whether the response expresses calibrated uncertainty or makes an unsupported diagnosis. Shaded rows denote Craft .
Model
Implementation
#Heads
Normal Lat.
Interv. Lat.
Δ Lat.
Normal FLOPs
Δ FLOPs
Hulu-Med 4B
Hooked masking
6
206.39 ms
210.31 ms
+1.90%
19539.96G
0.00%
Hulu-Med 4B
True head skipping
6
207.11 ms
208.10 ms
+0.48%
19539.96G
0.11% lower
InternVL3.5 4B
Hooked masking
10
43.27 ms
49.43 ms
+14.24%
3629.54G
0.00%
InternVL3.5 4B
True head skipping
10
43.91 ms
47.03 ms
+7.11%
3629.54G
0.10% lower
Table 14 : Online computational cost of head intervention on VQA-RAD. Hooked masking is the analysis implementation; true head skipping offers marginal FLOP savings.
Model
Setting
Loc. Samples
Total Forwards
Forwards / Sample
Param. Update
Hulu-Med 4B
VQA-RAD Before Q
391
111,342
284.8
✗
InternVL3.5 4B
VQA-RAD Before Q
698
158,446
227.0
✗
Qwen3-VL-8B
VQA-RAD Before Q
698
241,508
346.0
✗
Table 15 : Offline cost of head selection. The localization cost is paid once before deployment. Prompt forward counts include diagnostic passes for layer tracing, head scoring, and selection.
Figure 9 : Head-level tuned-lens preference distributions under textual conflict. Each panel shows one selected arbitration-sensitive head. The horizontal axis is the gold-versus-wrong preference margin, logp(y+)−logp(y−) , under the no-conflict condition (NC, green) and the injected-conflict condition (IC, red). Larger values indicate stronger preference for the gold answer. Across many heads, IC shifts left relative to NC, showing that textual conflict pushes head-level preferences from the correct answer toward the misleading answer.
Figure 10 : Head-level tuned-lens abstention distributions under visual degradation. Each panel shows one selected image-conflict-sensitive head. The horizontal axis is the abstention margin, logp(unknown)−logp(ybest) , where ybest is the best competing concrete answer. NC denotes the original image, while IC denotes the masked image with relevant evidence removed. If the model properly abstains after masking, IC should shift right. Instead, NC and IC largely overlap for many heads, indicating persistent concrete-answer commitment under degraded visual evidence.
Model (R / S)
Kpatch
Qwen3-4B
Llama3.2-3B
Hulu-Med-4B
Hulu-Med-7B
InternVL3.5-4B
Qwen3-VL-8B
2
0.9718 / 0.8577
0.9214 / 0.1475
0.9288 / 0.0671
0.9346 / 0.0648
0.9478 / 0.1633
0.9415 / 0.1586
4
0.9792 / 0.8553
0.9237 / 0.1483
0.9506 / 0.0687
0.9569 / 0.0665
0.9868 / 0.1775
0.9821 / 0.1694
8
0.9991 / 0.8449
0.9987 / 0.1501
0.9991 / 0.0714
0.9988 / 0.0698
0.9948 / 0.1781
0.9962 / 0.1726
16
0.9160 / 0.8983
0.9561 / 0.1459
0.9736 / 0.0741
0.9758 / 0.0726
0.9733 / 0.1818
0.9714 / 0.1769
32
0.8206 / 0.2564
0.9004 / 0.0442
0.8255 / 0.0153
0.8389 / 0.0147
0.8595 / 0.0535
0.8668 / 0.0512
Table 16 : Patch window sensitivity. Robustness (R) ↑ measures agreement with the local window consensus; Smoothness (S) ↓ measures adjacent layer variation. Best R and best S are bolded separately for each model.
Figure 11 : Downstream attention after arbitration head ablation. Heatmaps compare the same conflict input before and after excising the selected arbitration heads. Blue and red regions denote image tokens and injected conflict text, respectively. Ablation weakens downstream attention to the conflict text region, consistent with reduced textual override propagation.
β
# Selected Heads
ΔUR
HR
0.0
6
7.23
0.7477
0.2
7
9.68
0.7444
0.4
5
3.35
0.5581
0.6
6
3.23
0.4156
0.8
5
3.24
0.2455
1.0
4
3.23
-0.0001
Table 17 : Layer wise attenuation sensitivity. ΔUR denotes the post intervention increase in unknown rate on the held out image conflict split. HR is the layer level hallucination relief score in eq. 9 . The highlighted row denotes the setting used for brake tracing.
Noise std
#Heads
UR Pre
UR Post
ΔUR
24
6
1.82
35.45
33.64
36
6
3.18
36.14
32.95
48
6
7.05
40.45
33.41
64
6
11.59
43.64
32.05
Table 18 : Robustness to local visual noise. Brake head intervention remains effective under local corruption at varying noise strengths.
Model
Brake Heads
Arbitration Heads
Intersection
Union
Jaccard
Overlap Coef.
Hulu-Med 4B
6
10
0
16
0.000
0.000
InternVL3.5 4B
7
8
0
15
0.000
0.000
Table 19 : Head set overlap between arbitration and brake localization on SLAKE. Both Jaccard and overlap coefficient are zero, confirming disjoint head sets.
Figure 12 : Arbitration head distribution in InternVL3.5 and Qwen3-VL. Blue circles denote selected arbitration heads, green circles denote selected backbone heads, and pale dots denote unselected heads. Circle size is proportional to the corresponding score. Both models exhibit a broad arbitration band in mid to late layers, while backbone heads are more diffuse with a late-layer tail.
Model
Failure Mode
Transfer
#Heads
ΔResist
ΔUR
Hulu-Med 4B
Arbitration
SLAKE heads → VQA-RAD
10
30.77
–
Arbitration
VQA-RAD heads → SLAKE
6
12.44
–
Brake
HealMed heads → SLAKE
5
–
4.55
Brake
SLAKE heads → HealMed
6
–
21.56
InternVL3.5 4B
Arbitration
SLAKE heads → VQA-RAD
8
13.71
–
Arbitration
VQA-RAD heads → SLAKE
10
1.14
–
Table 20 : Cross dataset transfer results. Arbitration transfer reports ΔResist ; brake transfer reports ΔUR .
Benchmark
Transfer
#Heads
ΔResist
VQA-RAD
InternVL → Hulu-Med
10
17.95
VQA-RAD
Hulu-Med → InternVL
6
12.00
SLAKE VQA
InternVL → Hulu-Med
8
5.07
SLAKE VQA
Hulu-Med → InternVL
10
2.76
Table 21 : Cross model transfer of arbitration head sets. Transfer is reported as ΔResist .
Benchmark
Transfer
#Heads
ΔUR
SLAKE VQA
Hulu-Med → InternVL
6
0.00
SLAKE VQA
InternVL → Hulu-Med
7
-3.03
HealMed-VQA
Hulu-Med → InternVL
6
0.42
HealMed-VQA
InternVL → Hulu-Med
7
-2.76
Table 22 : Cross model transfer of brake head sets. Transfer is reported as ΔUR .
Figure 13 : Bootstrap visualisation of clean competence stratified head sets. Each point denotes a bootstrap sample of the head set under a given stratum. Shaded regions show kernel density estimates. The 100% and 80% strata remain nearby, whereas 50% and 30% strata shift toward an intermediate region and the all wrong stratum is clearly separated.
Model
Stratum
#Heads
Dominant layers
Overlap with C100
Interpretation
Hulu-Med 4B
100% clean correct
14
8, 12–15, 18–21
100%
Main clean conflict arbitration set.
80% clean correct
15
12–22, mainly 18–21
65–75%
Largely preserves the C100 core with neighbouring heads.
50% clean correct
21
14–27 plus 32–35
10–25%
Less localized; recruits difficulty sensitive heads.
30% clean correct
24
15–27 plus 32–35
5–20%
Shifts toward late answer commitment heads.
0% / all wrong
20
diffuse across 0–35
< 10%
No stable conflict core; resembles wrong prior heads.
InternVL3.5 4B
100% clean correct
21
13–21, 24, 34
100%
Main clean conflict arbitration set.
Table 23 : Clean competence stratified head set summary. The 100% stratum is the main setting. Lower strata are sensitivity checks only.
Train
Val
Position
Model
Method
CFR ↓
Resist ↑
C2W ↓
W2C ↑
CFR ↓
Resist ↑
C2W ↓
W2C ↑
Prefix
Qwen3-4B
Baseline
81.60
18.40
–
–
83.33
16.67
–
–
Random
91.04
8.96
11.79
2.36
94.25
5.75
12.07
1.15
Craft
33.49
66.51
0.47
48.58
37.36
62.64
0.00
45.98
Llama3.2-3B
Baseline
54.58
45.42
–
–
57.69
42.31
–
–
Random
38.33
61.67
27.08
43.33
42.31
57.69
26.92
42.31
Table 24 : Full train/val results on ConflictMedQA. Results across three injection positions. Shaded rows denote Craft .
Train
Val
Position
Model
Method
CFR ↓
Resist ↑
C2W ↓
W2C ↑
CFR ↓
Resist ↑
C2W ↓
W2C ↑
Prefix
Qwen3-4B
Baseline
48.40
51.60
–
–
49.67
50.33
–
–
Random
22.40
77.60
0.80
26.80
57.17
42.83
12.17
4.67
Craft
22.00
78.00
0.40
26.80
23.59
76.41
0.00
26.09
Llama3.2-3B
Baseline
56.80
43.20
–
–
58.46
41.54
–
–
Random
60.40
39.60
7.60
4.00
63.08
36.92
8.85
4.23
Table 25 : Full train/val results on PubMedQA. Results across three injection positions. Shaded rows denote Craft .
Train
Val
Position
Model
Method
CFR ↓
Resist ↑
C2W ↓
W2C ↑
CFR ↓
Resist ↑
C2W ↓
W2C ↑
Prefix
Hulu-Med 4B
Baseline
33.76
66.24
–
–
33.33
66.67
–
–
Random
34.53
65.47
2.81
2.05
33.59
66.41
1.79
1.54
Craft
16.11
83.89
0.77
18.41
15.38
84.62
0.26
18.21
Hulu-Med 7B
Baseline
26.40
73.60
–
–
25.80
74.20
–
–
Random
26.10
73.90
0.90
1.20
25.40
74.60
0.90
1.30
Table 26 : Full train/val results on VQA-RAD textual conflict. Results across three injection positions. Shaded rows denote Craft .
Train
Val
Position
Model
Method
CFR ↓
Resist ↑
C2W ↓
W2C ↑
CFR ↓
Resist ↑
C2W ↓
W2C ↑
Prefix
Hulu-Med 4B
Baseline
18.17
81.83
–
–
19.15
80.85
–
–
Random
21.32
78.68
3.93
0.70
21.28
78.72
2.95
0.82
Craft
11.16
88.84
3.86
10.87
12.11
87.89
5.07
12.11
Hulu-Med 7B
Baseline
16.00
84.00
–
–
15.40
84.60
–
–
Random
16.50
83.50
1.00
0.50
15.90
84.10
1.00
0.50
Table 27 : Full train/val results on SLAKE textual conflict. Results across three injection positions. Shaded rows denote Craft .
Train
Val
Benchmark
Model
Method
UR ↑
O2U ↑
U2O ↓
UR ↑
O2U ↑
U2O ↓
SLAKE
Hulu-Med 4B
Baseline
36.90
–
–
37.88
–
–
Random
44.23
8.10
0.77
45.45
8.33
0.76
Craft
64.20
27.30
0.00
65.15
27.27
0.00
Hulu-Med 7B
Baseline
43.40
–
–
44.12
–
–
Random
47.85
5.05
0.60
48.82
5.29
0.59
Table 28 : Full train/val results for brake failure under visual degradation. Craft ablates brake heads to recover abstention. ΔUR=O2U−U2O . Shaded rows denote Craft .
Figure 14 : Binary validation confusion matrices for arbitration intervention on Hulu-Med 4B and VQA-RAD.
Table 29 : Computational cost of head intervention on VQA-RAD. Hooked masking is the analysis implementation; true head skipping offers marginal FLOP savings.
Vision-language models (VLMs) generate fluent causal explanations, but current evaluations cannot distinguish linguistic plausibility from faithful causal reasoning. We introduce a dual-probe methodology that isolates these properties. The Text-Only Probe measures linguistic quality. The Chain-Text Probe requires models to first generate explicit causal chains. The Abstraction Gap (AG) metric quantifies the normalized performance difference. Evaluating eight VLMs on CAGE (Causal Abstraction Gap Evaluation), a benchmark of 49,500 questions across 5,500 images spanning Pearl's causal hierarchy, we find seven models exhibit AG exceeding 0.50 with text scores of 6--8 but chain scores below 2.5. Fine-tuning on 45,000 chain-annotated examples fails to close the gap. However, one model achieves near-zero AG. The capability exists within current VLM architectures and depends on pretraining and architectural choices. CAGE provides a diagnostic tool for assessing faithful causal reasoning in VLMs.
Chinh Hoang, Mohammad Rashedul Hasan
Department of Electrical and Computer Engineering, University of Nebraska–Lincoln, Lincoln, Nebraska, USA.
Vision-language models must reconcile visual evidence with memorized world knowledge when the two conflict. How they resolve this conflict shapes the reliability of multimodal systems, yet prior work characterizes it behaviorally without a component-level causal account. We combine activation patching across three granularities (residual stream, attention heads, and MLP sublayers) with model-component ablation studies and mechanistic analysis. Across three VLM families, we find that visual grounding emerges by default, whereas prior grounding depends on a small set of causally necessary attention heads (2.5-4.8%) concentrated in the second half of the network. These heads enable answers from stored world knowledge (e.g., "red" for a strawberry) despite conflicting visual input. Ablating them flips predictions from knowledge-grounded to visually grounded answers in 68-96% of cases under prior-knowledge prompts, but changes only 0.8-7.5% of visually grounded predictions, establishing an asymmetric causal structure. The identified heads decompose into routing heads, which modulate information flow, and writing heads, which directly project answer tokens into the residual stream. This structure is consistent across model families and scales, revealing a sparse causal circuit underlying perception-knowledge conflict in VLMs.
Vision-language models can describe an image with remarkable accuracy, yet a more fundamental question remains unanswered: what visual information actually drives their answers? In this work, we investigate this question through causal tracing, and we observe that highly causal vision tokens often lie outside the target region. Extending the analysis to larger vision-language models reveals a similar pattern across models and corruption settings, suggesting that strong multimodal performance does not necessarily imply spatially localized causal representations. We further investigate: can these models preserve visual structure when appearance cues are removed? and find that visual cues are exploited to understand visual structures. Together, our experiments expose a gap between seeing, using, and reasoning over visual structure, and provide a causal framework for studying how visual information is transformed, preserved, and ultimately used by modern vision-language models.
Naren Kumar S, Tirth Bhatt, Mayank Singh
LINGO Research Group, Indian Institute of Technology Gandhinagar, India