Visual Evidence Under Cross-Examination: Evaluating and Controlling Decision-Level Evidence Use in Vision-Language Models
Authors: Huiyao Zhang, Jin Bai, Zilong Su, Rui Guo, Chaofan Qin, Jinze Lv, Wenhui Yu, Hongfei Wang, +1 more
Organizations: Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences · University of Chinese Academy of Sciences · University of Science and Technology of China
Vision-language models increasingly reason through crops, regions, and tool-produced observations. Yet an observation can influence the answer without benefiting the candidate it supports. We study candidate-bound visual contribution: valid evidence should help, invalidating its supporting relation should remove its additional effect, and valid rebinding should redirect that effect to the newly supported candidate. We introduce CROSS-Bench, a benchmark of 28,000 decision problems, with matched invalidation and rebinding tests on a dedicated evaluation subset. Our RIVET interface preserves evidence identity and uncertainty, composes a candidate-conditioned response, and separately controls its strength. Shared-evidence experiments show that task accuracy and evidence ownership can diverge. Under matched capacity and training, RIVET increases normalized effect transfer from 0.512 to 0.651 where clean evidence has a positive effect. The advantage persists on common evaluation examples and across repeated decision-layer fits. With evidence predicted from raw inputs, RIVET improves CROSS-Bench accuracy by an average of 5.70 pp across four frozen backbones, relative to the same models without auxiliary evidence. These results separate the utility of visual evidence from the candidate-specific destination of its effect.
Figures & tables
Figure 1: Candidate-bound visual contribution. (A) The observation-to-candidate relation; (B) clean, invalid, and rebind request–evidence pairs; (C) utility–ownership dissociation under matched evidence on the 960-root matrix (Qwen2.5-VL-7B, predicted evidence). Ownership uses method-specific clean-active sets. Arrows denote schematic effect destinations, not magnitudes.
Figure 2: CROSS-Bench evaluation protocol. Controlled supplies evidence-access structure and Ground predicts it; both estimate visual state. Clean, invalid, and rebind conditions test utility, specificity, and ownership over 2–5 candidates.
Figure 3: Overview of RIVET. Specify plans visual requirements; Preserve acquires, binds, and reads observations through a shared typed interface. Compose uses candidate queries to organize and read the complete evidence set; Bound scales the resulting posterior response. Highlighted e∗ follows one illustrative support through the pipeline. Here qt=max(pt,ϵ) for t∈{0,1} .
Method
Acc. (%) ↑
Gclean↑
Ainv↓
Clean-active
Tauth↑/Lorig↓
(a) Released-system diagnosis
Qwen2.5-VL-7B (Base)
64.90
0.000
0.000
0/707
–
VLM-R 3 -7B
72.50
0.153
0.162
355/707
0.413 / 0.274
Pixel-Reasoner-7B
71.04
0.128
0.143
332/707
0.427 / 0.253
DeepEyes-7B
75.10
0.184
0.171
386/707
0.462 / 0.283
TreeVGR-7B
73.96
0.161
0.187
367/707
0.497 / 0.308
Table 1: CROSS-Bench results on the 960-root intervention matrix. (a) Released-system diagnosis; posterior metrics use each checkpoint’s no-auxiliary reference. (b) Shared predicted evidence and Qwen2.5-VL-7B Base. Clean-active counts give scoreable positive-effect roots out of 707 rebinding-eligible roots; ownership uses these method-specific sets (“–”: undefined).
Backbone
Base Acc. (%) ↑
Controlled Acc. (%) ↑
Ground Acc. (%) ↑
Ground Tauth↑/Lorig↓
Qwen2.5-VL-7B
66.34
73.18 (+6.84)
71.30 (+4.96)
0.651 / 0.205
Qwen3-VL-8B
67.96
75.89 (+7.93)
73.93 (+5.97)
0.638 / 0.216
InternVL3.5-8B
62.98
71.25 (+8.27)
69.45 (+6.47)
0.659 / 0.225
LLaVA-OneVision-2-8B
66.00
73.11 (+7.11)
71.39 (+5.39)
0.647 / 0.196
Mean over backbones
65.82
73.36 (+7.54)
71.52 (+5.70)
0.649 / 0.211
Table 2: Backbone reuse and task generalization. (a) CROSS-Bench accuracy on 5,600 roots (parentheses: gains over Base in pp); Ground ownership on the 960-root matrix with method-specific active sets. (b) Frozen Ground task generalization averaged over four backbones with shared predicted evidence.
Decision interface
Predicted state Tauth↑/Lorig↓
Oracle state Tauth↑/Lorig↓
Unconstrained Fusion
0.072 / 0.657
0.101 / 0.612
Capacity-matched Fusion
0.512 / 0.286
0.564 / 0.253
RIVET
0.651 / 0.205
0.685 / 0.171
Table 3: State-only Oracle and architectural controls. (a) Oracle retains predicted bindings, active sets, and clean-effect normalizers per interface; owner labels remain hidden. (b) Qwen2.5-VL-7B Ground: accuracy on 5,600 roots (parentheses: gains over Base in pp); relation metrics on the 960-root matrix, with ownership fixed to Full RIVET’s 521 active roots and clean-effect normalizers.
Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer causally depends on the local evidence that supports it. We propose Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding. For each response, CED neutralizes an object-centric Evidence Region and compares the resulting support drop against matched non-evidence Regions. We combine this signal with answer correctness inside GRPO, rewarding correct answers that rely on the evidence path rather than shortcut or nuisance paths. CED uses weak object-level proposals, requires no question-specific evidence annotations, and adds no inference-time overhead. Across nine public benchmarks and four backbones, CED outperforms prior RL-based post-training methods, with targeted analyses verifying its object-centric signal.
Haojie Huang, Xinlei Yu, Chengming Xu +6
National University of Singapore · Zhejiang University · Fudan University +2
Multimodal large language models (MLLMs) fail at fine-grained visual questions less because they cannot reason than because they never see the evidence: high-resolution images are downsampled before encoding, so the model answers from linguistic priors. The standard remedies are expensive: annotated answers (SFT), hand-engineered verifiers (RLVR), or a large external teacher (on-policy distillation). We ask whether the visual evidence itself can supply the signal for free. We formalize the contrastive evidence gap, the per-token log-likelihood ratio that a model assigns to its own output when conditioned on a question-relevant region versus an irrelevant one, and study it across Qwen2.5-VL-7B, Qwen3-VL-8B, and Qwen3-VL-30B-A3B on V*Bench. Our main positive result is training-free: selecting the candidate crop under which the model's answer distribution is most peaked, using a single-view, label-free criterion, discovers the answer-bearing region with no bounding boxes, training, or labels. It localizes the target 4.4 to 5.1 times better than chance and raises fine-grained accuracy from 70 percent to 85 percent at inference. We further show that the gap is complementary to the model's own confidence. Combining them predicts correctness better than either alone, with AUC up to 0.99, and flags confidently wrong answers, with AUC ranging from 0.97 to 1.00 within the high-confidence subset. All effects concentrate on perception-bottleneck questions and vanish on a global-context control. Finally, we report an honest negative result: converting the same signal into a training method, gated self-distillation (SEG-Distill), does not outperform the base model at pilot scale across three gate designs, while more aggressive gating degrades accuracy. The signal is real, but converting it into training gains remains an open problem.
Santi Ram Tiwari, Nihal Naik, Devbrat Pandey +1
KGraph AI Solutions Pvt. Ltd. Bangalore, India – 560016
When a multimodal large language model answers a visual reasoning question correctly, is the prediction actually supported by the task-critical visual evidence? Correct answers can coexist with flawed reasoning, making accuracy alone an incomplete test of grounding. We introduce VisualFLIP, a paired benchmark with 1,374 images arranged as same-question perturbation pairs across cardinality, attribute, spatial, and logic tasks. Each pair keeps the question fixed but minimally changes the evidence so the gold answer deterministically flips. We evaluate 24 MLLMs with pair accuracy, which requires solving both sides of a pair, and Collapse Rate (CR), which measures how often a model that solves at least one side repeats the same non-empty answer for both images. Together, these metrics show that paired correctness and evidence dependence are related but distinct: capable models can still fail to update after task-critical visual changes, and collapse becomes more severe for some models when the edited image follows an earlier answer in a sequential setting. Further details are available on our project page: https://didizhu-judy.github.io/VisualFLIP/