Visual Evidence Under Cross-Examination: Evaluating and Controlling Decision-Level Evidence Use in Vision-Language Models
Authors: Huiyao Zhang, Jin Bai, Zilong Su, Rui Guo, Chaofan Qin, Jinze Lv, Wenhui Yu, Hongfei Wang, +1 more
Organizations: Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences · University of Chinese Academy of Sciences · University of Science and Technology of China
Vision-language models increasingly reason through crops, regions, and tool-produced observations. Yet an observation can influence the answer without benefiting the candidate it supports. We study candidate-bound visual contribution: valid evidence should help, invalidating its supporting relation should remove its additional effect, and valid rebinding should redirect that effect to the newly supported candidate. We introduce CROSS-Bench, a benchmark of 28,000 decision problems, with matched invalidation and rebinding tests on a dedicated evaluation subset. Our RIVET interface preserves evidence identity and uncertainty, composes a candidate-conditioned response, and separately controls its strength. Shared-evidence experiments show that task accuracy and evidence ownership can diverge. Under matched capacity and training, RIVET increases normalized effect transfer from 0.512 to 0.651 where clean evidence has a positive effect. The advantage persists on common evaluation examples and across repeated decision-layer fits. With evidence predicted from raw inputs, RIVET improves CROSS-Bench accuracy by an average of 5.70 pp across four frozen backbones, relative to the same models without auxiliary evidence. These results separate the utility of visual evidence from the candidate-specific destination of its effect.
Figures & tables
Figure 1: Candidate-bound visual contribution. (A) The observation-to-candidate relation; (B) clean, invalid, and rebind request–evidence pairs; (C) utility–ownership dissociation under matched evidence on the 960-root matrix (Qwen2.5-VL-7B, predicted evidence). Ownership uses method-specific clean-active sets. Arrows denote schematic effect destinations, not magnitudes.
Figure 2: CROSS-Bench evaluation protocol. Controlled supplies evidence-access structure and Ground predicts it; both estimate visual state. Clean, invalid, and rebind conditions test utility, specificity, and ownership over 2–5 candidates.
Figure 3: Overview of RIVET. Specify plans visual requirements; Preserve acquires, binds, and reads observations through a shared typed interface. Compose uses candidate queries to organize and read the complete evidence set; Bound scales the resulting posterior response. Highlighted e∗ follows one illustrative support through the pipeline. Here qt=max(pt,ϵ) for t∈{0,1} .
Method
Acc. (%) ↑
Gclean↑
Ainv↓
Clean-active
Tauth↑/Lorig↓
(a) Released-system diagnosis
Qwen2.5-VL-7B (Base)
64.90
0.000
0.000
0/707
–
VLM-R 3 -7B
72.50
0.153
0.162
355/707
0.413 / 0.274
Pixel-Reasoner-7B
71.04
0.128
0.143
332/707
0.427 / 0.253
DeepEyes-7B
75.10
0.184
0.171
386/707
0.462 / 0.283
TreeVGR-7B
73.96
0.161
0.187
367/707
0.497 / 0.308
Table 1: CROSS-Bench results on the 960-root intervention matrix. (a) Released-system diagnosis; posterior metrics use each checkpoint’s no-auxiliary reference. (b) Shared predicted evidence and Qwen2.5-VL-7B Base. Clean-active counts give scoreable positive-effect roots out of 707 rebinding-eligible roots; ownership uses these method-specific sets (“–”: undefined).
Backbone
Base Acc. (%) ↑
Controlled Acc. (%) ↑
Ground Acc. (%) ↑
Ground Tauth↑/Lorig↓
Qwen2.5-VL-7B
66.34
73.18 (+6.84)
71.30 (+4.96)
0.651 / 0.205
Qwen3-VL-8B
67.96
75.89 (+7.93)
73.93 (+5.97)
0.638 / 0.216
InternVL3.5-8B
62.98
71.25 (+8.27)
69.45 (+6.47)
0.659 / 0.225
LLaVA-OneVision-2-8B
66.00
73.11 (+7.11)
71.39 (+5.39)
0.647 / 0.196
Mean over backbones
65.82
73.36 (+7.54)
71.52 (+5.70)
0.649 / 0.211
Table 2: Backbone reuse and task generalization. (a) CROSS-Bench accuracy on 5,600 roots (parentheses: gains over Base in pp); Ground ownership on the 960-root matrix with method-specific active sets. (b) Frozen Ground task generalization averaged over four backbones with shared predicted evidence.
Decision interface
Predicted state Tauth↑/Lorig↓
Oracle state Tauth↑/Lorig↓
Unconstrained Fusion
0.072 / 0.657
0.101 / 0.612
Capacity-matched Fusion
0.512 / 0.286
0.564 / 0.253
RIVET
0.651 / 0.205
0.685 / 0.171
Table 3: State-only Oracle and architectural controls. (a) Oracle retains predicted bindings, active sets, and clean-effect normalizers per interface; owner labels remain hidden. (b) Qwen2.5-VL-7B Ground: accuracy on 5,600 roots (parentheses: gains over Base in pp); relation metrics on the 960-root matrix, with ownership fixed to Full RIVET’s 521 active roots and clean-effect normalizers.