Large vision-language models (LVLMs) exhibit strong reasoning capabilities, yet the visual and textual evidence supporting the generated responses remains difficult to identify. Faithful token attribution explains an LVLM's response by assigning scores that rank image and prompt tokens by how much the model relies on them, such that removing higher-ranked tokens causes the likelihood of the generated response to drop more rapidly. However, existing token-attribution methods have been developed mainly for text-based language models, and our empirical study reveals two challenges when complex multimodal sources are involved. First, the joint image-text attribution can underrepresent visual evidence relative to text, obscuring the image regions supporting the response. Second, visual evidence may influence the generated response through multiple intermediate reasoning paths, while existing methods trace only a limited subset of these paths, causing important visual contributions to be underestimated. Motivated by these insights, we introduce VTrace, a multimodal token-attribution framework that traces input contributions through intermediate reasoning and calibrates attribution scores across modalities. VTrace constructs pairwise attributions that highlight token-specific contributions and aggregates all forward attribution paths in closed form to account for both direct and indirect contributions. Cross-modal calibration then rescales image and text attribution scores using modality contributions estimated from response-likelihood changes, enabling a unified ranking of input tokens. Evaluations against seven baselines across six visual reasoning benchmarks demonstrate the superior attribution faithfulness. Project page: https://vtrace-attribution.github.io/.
Figures & tables
Figure 2: Tracing only highly-attributed paths misses evidence that many weaker paths carry.
Figure 3: Existing methods overlook visual cues.
Figure 4: Overview of VTrace . VTrace (1) measures each pair of tokens against the average contribution of its own modality and (2) aggregates all direct and indirect paths in closed form, (3) enabling image patches and words to be faithfully ranked together.
Dataset
Setting
RISE
ReAGent
HETA
FlowTracer
IFR
Attn Rollout
AttnLRP
FlashTrace
VTrace
MMStar
Image
Ins. ↑
0.505
0.497
0.532
0.539
0.539
0.563
0.555
0.600
Del. ↓
0.458
0.463
0.412
0.409
0.439
0.384
0.394
0.352
Joint
Ins. ↑
0.371
0.489
0.506
0.507
0.501
0.529
0.554
0.581
Del. ↓
0.325
0.245
0.226
0.222
0.273
0.219
0.204
0.180
MathVista
Image
Ins. ↑
0.500
0.514
0.579
0.591
0.584
0.599
0.600
0.662
Del. ↓
0.447
0.444
0.368
0.363
0.394
0.358
0.354
0.311
Table 1: RISE attribution faithfulness on Qwen3-VL-8B across six benchmarks under Image and Joint perturbation settings. We report RISE insertion ( Ins. ↑ , higher is better) and deletion ( Del. ↓ , lower is better) AUC. The Image variant perturbs image patch tokens only, while Joint perturbs both image and text tokens. Best and Runner-up are highlighted.
Method
Image
Joint
Ins. ↑
Del. ↓
Ins. ↑
Del. ↓
VTrace
0.631
0.327
0.610
0.187
w/o centering
0.625
0.340
0.599
0.200
w/o multi-hop
0.578
0.362
0.470
0.273
w/o calibration
0.631
0.327
0.602
0.200
Table 2: Ablation study on different components
Figure 5: Attribution faithfulness across different LVLM sizes and families.
Correct Predictions
Incorrect Predictions
Variant
Method
Del. RISE ↓
Del. MAS ↓
Ins. RISE ↑
Ins. MAS ↑
Del. RISE ↓
Del. MAS ↓
Ins. RISE ↑
Ins. MAS ↑
Image
Best Baseline
0.377
0.508
0.590
0.451
0.424
0.579
0.603
0.460
VTrace
0.333
0.468
0.644
0.514
0.398
0.550
0.636
0.496
Joint
Best Baseline
0.209
0.338
0.581
0.381
0.200
0.323
0.605
0.410
VTrace
0.171
0.235
0.626
0.463
0.176
0.241
0.640
0.474
Table 3: Attribution faithfulness on correct and incorrect model predictions over benchmarks. We compare against the strongest baseline for each metric independently.
Method
MMStar
MathVista
MMMU
MathVerse
MMMU-Pro
HallusionBench
RealWorldQA
MathVision
Avg.
Base
63.53
73.30
55.33
56.55
47.57
70.87
72.55
41.78
60.19
GRPO
69.53
77.40
61.89
62.64
51.68
72.77
72.29
43.75
63.99
+ VTrace
70.20
77.20
65.00
65.13
51.45
73.29
73.20
46.71
65.27
Table 4: Attribution-guided learning with VTRACE on Qwen3-VL-4B. Incorporating VTrace attribution into GRPO improves the average performance across benchmarks.
Figure 8: Qualitative study: Attribution comparison between diverse hops and IFR.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Metric
ReAGent
HETA
FlowTracer
IFR
Attn Rollout
AttnLRP
FlashTrace
VTrace
MMStar
RISE ins↑
0.505
0.497
0.532
0.539
0.539
0.563
0.555
0.600
MAS ins↑
0.336
0.329
0.368
0.384
0.399
0.392
0.410
0.460
RISE del↓
0.458
0.463
0.412
0.409
0.439
0.384
0.394
0.352
MAS del↓
0.620
0.625
0.563
0.551
0.566
0.547
0.528
0.489
MathVista
RISE ins↑
0.500
0.514
0.579
0.591
0.584
0.599
0.600
0.662
MAS ins↑
0.334
0.340
0.421
0.449
0.452
0.427
0.464
0.535
Appendix
Table 5: Attribution faithfulness of the Image variant on Qwen3-VL-8B across six benchmarks. We report RISE and MAS insertion ( ins↑ , higher is better) and deletion ( del↓ , lower is better) AUC and misalignment scores. The Image variant perturbs image patch tokens only. Best results are highlighted in red bold , and second-best results are highlighted in blue underlining .
Dataset
Metric
ReAGent
HETA
FlowTracer
IFR
Attn Rollout
AttnLRP
FlashTrace
VTrace
MMStar
RISE ins↑
0.371
0.489
0.506
0.507
0.501
0.529
0.554
0.581
MAS ins↑
0.165
0.277
0.299
0.305
0.314
0.314
0.349
0.397
RISE del↓
0.325
0.245
0.226
0.222
0.273
0.219
0.204
0.180
MAS del↓
0.494
0.408
0.372
0.364
0.384
0.341
0.332
0.246
MathVista
RISE ins↑
0.382
0.467
0.517
0.516
0.501
0.512
0.574
0.647
MAS ins↑
0.186
0.268
0.316
0.324
0.327
0.304
0.377
0.497
Appendix
Table 6: Attribution faithfulness of the Joint variant on Qwen3-VL-8B across six benchmarks. The Joint variant perturbs both image and text tokens.
Figure 9: Sensitivity analysis for VTrace to the decay rate γ .
Figure 10: Sensitivity analysis for IFR to the decay rate γ .
Dataset
Metric
ReAGent
HETA
FlowTracer
IFR
Attn Rollout
AttnLRP
FlashTrace
VTrace
Image variant
MMStar
RISE ins↑
0.531
0.531
0.531
0.525
0.555
0.575
0.539
0.614
MAS ins↑
0.382
0.364
0.368
0.376
0.421
0.419
0.399
0.482
RISE del↓
0.480
0.453
0.454
0.462
0.452
0.418
0.447
0.386
MAS del↓
0.631
0.620
0.614
0.609
0.586
0.577
0.587
0.531
MathVerse
RISE ins↑
0.490
0.600
0.612
0.611
0.618
0.596
0.617
0.673
Appendix
Table 7: Attribution faithfulness on Qwen3-VL-4B .
Dataset
Metric
ReAGent
HETA
FlowTracer
IFR
Attn Rollout
AttnLRP
FlashTrace
VTrace
Image variant
MMStar
RISE ins↑
0.523
0.639
0.592
0.614
0.659
0.618
0.635
0.694
MAS ins↑
0.348
0.492
0.459
0.486
0.554
0.477
0.520
0.588
RISE del↓
0.478
0.381
0.411
0.388
0.387
0.384
0.372
0.327
MAS del↓
0.650
0.533
0.547
0.524
0.510
0.532
0.499
0.461
MathVerse
RISE ins↑
0.405
0.653
0.623
0.659
0.680
0.616
0.672
0.709
Appendix
Table 8: Attribution faithfulness on InternVL3.5-2B .
Dataset
Metric
ReAGent
HETA
FlowTracer
IFR
Attn Rollout
AttnLRP
FlashTrace
VTrace
Image variant
MMStar
RISE ins↑
0.529
0.684
0.637
0.679
0.681
0.646
0.694
0.726
MAS ins↑
0.364
0.570
0.506
0.572
0.580
0.517
0.592
0.633
RISE del↓
0.479
0.344
0.365
0.338
0.370
0.367
0.325
0.303
MAS del↓
0.643
0.475
0.505
0.462
0.491
0.509
0.447
0.426
MathVerse
RISE ins↑
0.421
0.718
0.695
0.732
0.733
0.662
0.742
0.753
Appendix
Table 9: Attribution faithfulness on InternVL3.5-8B .
Dataset
Split
Metric
ReAGent
HETA
FlowTracer
IFR
Attn Rollout
AttnLRP
FlashTrace
VTrace
MMStar
Correct
RISE ins↑
0.371
0.482
0.504
0.504
0.497
0.521
0.552
0.581
MAS ins↑
0.166
0.271
0.297
0.303
0.310
0.306
0.347
0.400
RISE del↓
0.326
0.251
0.230
0.225
0.278
0.225
0.207
0.181
MAS del↓
0.494
0.417
0.377
0.369
0.390
0.351
0.336
0.248
Incorrect
RISE ins↑
0.371
0.511
0.514
0.514
0.513
0.551
0.559
0.580
MAS ins↑
0.164
0.295
0.305
0.311
0.324
0.337
0.352
0.389
Appendix
Table 10: Attribution faithfulness on correctly and incorrectly answered samples for the Joint variant . RISE and MAS are reported under insertion ( ins↑ , higher is better) and deletion ( del↓ , lower is better). Best results are highlighted in red bold , and second-best results are highlighted in blue underlining .
Dataset
Split
Metric
ReAGent
HETA
FlowTracer
IFR
Attn Rollout
AttnLRP
FlashTrace
VTrace
MMStar
Correct
RISE ins↑
0.499
0.493
0.534
0.542
0.541
0.560
0.556
0.603
MAS ins↑
0.328
0.326
0.371
0.390
0.403
0.389
0.414
0.466
RISE del↓
0.453
0.459
0.402
0.399
0.429
0.380
0.386
0.344
MAS del↓
0.616
0.618
0.549
0.536
0.552
0.539
0.514
0.478
Incorrect
RISE ins↑
0.525
0.509
0.526
0.531
0.531
0.572
0.550
0.591
MAS ins↑
0.357
0.336
0.360
0.369
0.388
0.399
0.399
0.442
Appendix
Table 11: Attribution faithfulness on correctly and incorrectly answered samples for the Image variant . RISE and MAS are reported under insertion ( ins↑ , higher is better) and deletion ( del↓ , lower is better). Best results are highlighted in red bold , and second-best results are highlighted in blue underlining .
Figure 11: Fine-grained insertion and deletion perturbation curves across benchmarks.
Figure 12: (a) How far back a generated token’s credit comes from. (b) The answer’s image credit as hops are summed. (c) Performance comparison between diverse hops, VTrace ( R ) and IFR.
Figure 13: Image token attribution on four MMStar questions. Each method’s top 20% of patches is outlined, with the share of the measured dependence it captures in the corner (higher is better). The measured dependence is the likelihood drop from blurring each patch in turn. VTRACE leads on the first two questions, ties FlowTracer on the third, and trails AttnLRP on the fourth.
Figure 14: Recognizing the sport. The model answers “Soccer”. VTrace places most of its top patches on the goal net, the object that identifies the sport, while IFR and FlowTracer spread more patches over the players and the grass. In the text, VTrace scores the option “Soccer” and the concluding “soccer” highest, and also shades the scene the model describes, such as “large net”, “goal frame” and “soccer ball”; FlowTracer concentrates on the question wording.
Figure 15: Reading the side lengths of a parallelogram. The model adds the side lengths shown in the image, 2×23+2×16=78 , and answers (D). VTrace ’s top patches cover the “23 ft” and “16 ft” labels, while IFR’s lie mostly along the image border. In the text, IFR highlights little beyond the answer options, whereas VTrace highlights “perimeter” in the question and the side lengths and sums in the reasoning that lead to the answer.
Figure 16: Reading values from a table. The model reads that the count fell from 23 in 2014 to 22 in 2015 and answers (A) −1 . VTrace ’s top patches fall on the year column and the 2014 and 2015 rows, while Attention Rollout concentrates on the table’s title bar. In the text, Attention Rollout highlights setup words such as “Bloomington Consulting” and “calculate”, whereas VTrace highlights the years, the change in employees and the result −1 .
Figure 17: A harder case: the ball closest to the man. The model answers (A) White. In the text, VTrace highlights the key question words “ball closest” and “man” and the reasoning about the white cue ball. In the image, its top patches cover the balls on the table, including the white cue ball and the red ball at the man’s hand, but many also fall on the dark background above the table.
Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains difficult to inspect. Recent logit-lens attribution methods project each visual-token hidden state into the vocabulary space to explain generated words, but this token-wise readout introduces a mismatch: visual tokens are context-mixed by the model, while the attribution score is decoded independently at each token location. This often produces fragmented attribution maps and can be further affected by autoregressive context signals from preceding text tokens. We propose ERCR, an attribution framework built from Evidence Recomposition (ER) and Predictive Context Residualization (PCR). ER aggregates target evidence across multiple views with different token-to-region assignments, reducing attribution fragmentation caused by a single readout grid. PCR estimates a preceding-token context map with RBO-based rank relevance and subtracts its fitted component from the ER map to suppress context-token interference. Experiments on LLaVA, Qwen2-VL, and InternVL families across COCO Caption, GranDf, and OpenPSG show that ERCR improves visual evidence for target tokens and mitigates preceding-token context interference under the existing evaluation protocol. On Qwen2-VL-2B, ERCR improves TAM F1-IoU from 39.10 to 44.45 on COCO Caption and from 30.83 to 37.20 on GranDf. Overall, ERCR provides a practical refinement for token-level visual evidence inspection.
Jiawei Liang, Jianjie Huang, Ruoyu Chen +4
Shenzhen Campus of Sun Yat-sen University, Shenzhen 518107, China · Zhongguan-cun Academy, Beijing 100094, China · University of Chinese Academy of Sciences, Beijing 100049, China +2
Vision-language models can describe an image with remarkable accuracy, yet a more fundamental question remains unanswered: what visual information actually drives their answers? In this work, we investigate this question through causal tracing, and we observe that highly causal vision tokens often lie outside the target region. Extending the analysis to larger vision-language models reveals a similar pattern across models and corruption settings, suggesting that strong multimodal performance does not necessarily imply spatially localized causal representations. We further investigate: can these models preserve visual structure when appearance cues are removed? and find that visual cues are exploited to understand visual structures. Together, our experiments expose a gap between seeing, using, and reasoning over visual structure, and provide a causal framework for studying how visual information is transformed, preserved, and ultimately used by modern vision-language models.
Naren Kumar S, Tirth Bhatt, Mayank Singh
LINGO Research Group, Indian Institute of Technology Gandhinagar, India
Large Vision-Language Models (LVLMs) have shown strong promise for multimodal reasoning, yet often struggle with tasks requiring concepts beyond what is directly observable in the input image. Existing methods generate intermediate images or latent visual tokens to guide reasoning, but these representations can introduce errors and increasingly interfere with textual reasoning as reasoning progresses. We propose Visual-to-Text Chain-of-Thought Distillation (V2T), a framework that enables LVLMs to internalize visual reasoning without generating intermediate visual representations at inference time. V2T first trains a teacher LVLM using interleaved visual and textual chains of thought, and then uses knowledge distillation to train a student LVLM using the teacher's logits and cross-entropy supervision from ground-truth textual reasoning. When reasoning images can be mapped to the original image, V2T can additionally distill the teacher's attention to corresponding regions, while ground-truth bounding boxes can further guide a subsequent reinforcement learning stage. Experiments across multiple multimodal reasoning benchmarks show that V2T consistently outperforms the teacher and existing baselines, improving average accuracy by 14.3% on a held-out set and 2.7% on the broader visual evaluation suite. Moreover, lightweight SFT and substantially reduced RL make V2T up to 42x faster to train than state-of-the-art baselines.
Wenhan Yang, Nilay Naharas, Ali Payani +1
University of California, Los Angeles · Cisco Research