CCRV-Bench: Constraint-Based Evaluation of Causal Reasoning in Vision-Language Models
Authors: Linyuan Gao, Yuan Wu, Yi Chang
Organizations: School of Artificial Intelligence, Jilin University Changchun 130012, China · Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, Jilin University Changchun 130012, China · International Center of Future Science, Jilin University Changchun 130012, China
Vision-language models (VLMs) have demonstrated excellent performance in visual tasks, but their visual causal reasoning capabilities still lack reliable evaluation. Existing evaluations struggle to distinguish whether a model is performing causal reasoning based on visual evidence or relying on statistical correlations for shortcut learning, thereby potentially overestimating their actual capabilities. This paper proposes CCRV-Bench, a constraint-driven visual causal reasoning benchmark for single-image physical scenarios. We construct an orthogonal framework that evaluates four causal task dimensions: causal relation discovery, state prediction, causal diagnosis, and intervention-outcome prediction. We further introduce entity symbolization, spatial grounding, the factual adversarial constraint, and minimalist output constraints to reduce shortcut cues while preserving the physical commonsense required by the task. Experiments across 14 multimodal models show that constraint sensitivity is task- and model-dependent: intervention-outcome prediction has the largest average effective degradation among the four causal tasks, spatial grounding is the most damaging constraint on average, and the factual adversarial constraint improves DCR for all evaluated models. These results show that unconstrained performance does not determine constrained robustness and that a single aggregate score can obscure distinct failures in causal identification, spatial grounding, and constraint-compliant expression. CCRV-Bench provides a standardized framework for diagnosing image-grounded causal reasoning under controlled constraints. The code is available at https://github.com/0815linyuan/CCRVBench
Figures & tables
Benchmark
Causal Capability
Constraint
Input
Process
Und.
Pred.
Attr.
Interv.
Ground.
Factual Adversarial Constraint
Shortcut
Modality
Proc. Eval
CELLO Chen et al. (2024)
✓
✓
✗
✓
✗
✗
✗
Image
✗
CausalVLBench Komanduri et al. (2025)
✗
✓
✓
✓
✗
✗
✗
Image
✗
CausalVQA Foss et al. (2025)
✓
✓
✓
✓
✗
✗
✗
video
✗
GQA Hudson and Manning (2019)
✗
✗
✓
✗
✗
✗
✗
Image
✗
CFBench Zhang et al. (2025)
✗
✗
✗
✗
✗
✗
✓
Text
✗
Table 1: Comparison of CCRV-Bench and related benchmarks across causal tasks, constraint mechanisms, input modalities, and evaluation granularity. “Und.”, “Pred.”, “Attr.”, and “Interv.” denote causal structure understanding, prediction, attribution, and intervention-outcome prediction, respectively. “Ground.” indicates explicit spatial grounding requirements; “Factual Adversarial Constraint” indicates verification of false textual premises against visual evidence; and “Shortcut” indicates controls targeting shortcut cues. “Modality” specifies the input modality, while “Proc. Eval” indicates evaluation of intermediate reasoning components beyond the final answer. Checkmarks indicate explicitly evaluated capabilities or implemented mechanisms.
Figure 1: Overview of CCRV-Bench. The top panel presents the data construction pipeline, yielding 800 baseline QA pairs and 3,200 constrained variants. The lower-left panel organizes four causal tasks and four constraint mechanisms into 16 evaluation settings. The right panels show model response generation, evaluation, and the reported metrics.
Evaluated Models
DCRbaseline
DCRconstraint
CSR
eDCR
CDI
eCDI
L1
L2
L3
DCR
L1
L2
L3
DCR
GPT-5.5
92.6
89.3
86.1
89.3
86.9
79.2
70.8
79.0
95.0
76.9
10.4
12.4
Claude-Opus-4.7
95.8
91.0
85.3
90.7
84.7
76.8
70.2
77.2
93.0
74.5
13.5
16.1
Doubao-Seed-2-Pro
95.4
91.9
82.3
89.8
70.2
61.6
55.8
62.6
91.9
60.0
27.3
29.8
Gemini-3.1-Pro-Preview
93.9
88.3
79.6
87.3
71.9
59.6
48.5
60.0
92.4
58.6
27.3
28.6
Gemini-3.5-Flash
94.9
90.5
84.0
89.8
69.2
59.8
51.3
60.1
86.3
56.2
29.7
33.6
Table 2: Main evaluation results of 14 VLMs on CCRV-Bench. L1–L3 report cascaded pass rates, and DCR is their mean. DCR, CSR, and eDCR are reported as percentages; CDI and eCDI are reported in percentage points. The best and second-best scores are highlighted in bold and underlined, respectively.
Evaluated Models
DCRbaseline
DCRconstraint
eCDI
X1
X2
X3
X4
X1
X2
X3
X4
X1
X2
X3
X4
GPT-5.5
82.3
92.7
88.2
94.2
82.3
78.4
78.6
76.6
1.8
16.5
11.8
19.5
Claude-Opus-4.7
89.7
91.0
89.3
92.7
77.0
76.8
76.9
78.1
15.0
17.0
15.8
16.7
Doubao-Seed-2-Pro
91.5
90.0
87.3
90.5
67.1
61.3
63.8
58.0
27.3
30.8
27.4
33.7
Gemini-3.1-Pro-Preview
85.7
85.8
88.0
89.5
65.4
56.6
60.4
57.5
22.0
29.8
29.8
33.0
Gemini-3.5-Flash
86.0
88.5
91.7
93.0
64.3
58.3
60.9
56.9
25.6
34.5
34.4
40.1
Table 3: Detailed cognitive performance breakdown across four causal task dimensions ( X1 : Discovery, X2 : Prediction, X3 : Diagnosis, X4 : Intervention-Outcome Prediction). DCR values are percentages and eCDI values are percentage-point drops (the internal differences are multiplied by 100). Best results are bolded , and second-best are underlined .
Evaluated Models
DCRbaseline
DCRconstraint
eCDI
Y1
Y2
Y3
Y4
Y1
Y2
Y3
Y4
GPT-5.5
89.3
77.0
75.5
97.7
65.7
12.5
21.5
-8.4
23.9
Claude-Opus-4.7
90.7
75.5
74.8
95.2
63.3
17.2
23.8
-4.2
27.8
Doubao-Seed-2-Pro
89.8
72.4
32.4
92.3
53.1
22.7
62.2
-2.5
36.7
Gemini-3.1-Pro-Preview
87.3
48.3
32.3
94.7
64.8
39.0
60.1
-7.0
22.5
Gemini-3.5-Flash
89.8
48.6
37.6
95.3
58.8
47.0
61.7
-5.2
31.0
Table 4: Detailed performance breakdown across four specific adversarial constraints ( Y1 : Entity Symbolization, Y2 : Entity-Region Spatial Grounding, Y3 : Factual Adversarial Constraint, Y4 : Minimalist Output). DCR values are percentages and eCDI values are percentage-point drops (the internal differences are multiplied by 100). Best results are bolded , and second-best are underlined .
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Cascaded pass rates for entity grounding (L1), causal mechanism (L2), and task-specific causal conclusion (L3) across 14 models under baseline and constrained settings. Each stacked bar comprises the complete-chain pass rate (L3), the L2–L3 gap, and the L1–L2 gap; its total height equals the L1 pass rate.
Model
Vanilla DCR
CoT-A DCR
CoT-B DCR
Y3 eDCR
Claude-Opus-4.7
90.7
89.6
88.5
94.9
Doubao-Seed-2-Pro
89.8
88.9
80.8
92.3
GPT-5.5
89.3
88.6
89.5
97.7
GLM-5V-Turbo
77.5
68.9
84.8
92.8
Appendix
Table 5: Comparison with general-purpose reasoning prompts. All values are percentages; the first three score columns report DCR, and the final column reports effective DCR under Y3 .
Model
DCR (Image/Text-only)
L1 raw (Image/Text-only)
L2 raw (Image/Text-only)
L3 raw (Image/Text-only)
Claude-Opus-4.7
90.67/42.17
95.75/48.13
92.75/49.75
89.00/69.75
Gemini-3.1-Pro-Preview
87.25/14.50
93.88/19.63
92.13/25.25
86.25/51.25
GPT-5.5
89.33/30.37
92.63/36.75
91.88/38.75
92.63/64.75
InternVL3-38B
68.13/18.00
79.38/25.63
73.13/23.13
72.38/46.88
Qwen3.6-35B-A3B
72.54/8.33
80.50/11.00
78.38/17.38
80.38/47.88
Average
81.58/22.67
88.42/28.23
85.65/30.85
84.12/56.10
Appendix
Table 6: Text-only diagnostic control. Entries report image-based/text-only scores in percent; L1–L3 are raw binary rates, whereas DCR uses the cascaded scoring rule. Averages are computed from unrounded scores.
Model
Baseline
Text-only+ Y1
Oracle-informed+ Y1
Gain over Text-only (points)
Claude-Opus-4.7
90.67
71.87
97.79
+25.92
Gemini-3.1-Pro-Preview
87.25
58.21
85.42
+27.21
GPT-5.5
89.33
62.29
86.08
+23.79
InternVL3-38B
68.13
45.24
74.92
+29.67
Qwen3.6-35B-A3B
72.54
36.50
74.83
+38.34
Average
81.58
54.82
83.81
+28.99
Appendix
Table 7: Image-based baseline and no-image Y1 controls (DCR in percent). The Oracle-informed condition supplies the symbolized adapted reference answer. It is an answer-informed diagnostic control, not a strict edge-only intervention. The gain is Oracle-informed minus Text-only. Averages and gains are computed from unrounded scores.
Model
Baseline
Y1 : constrained → Oracle-informed
Y2 : constrained → Oracle-informed
Y4 : constrained → Oracle-informed
GPT-5.5
89.33
77.04 → 80.62 (+3.58)
75.46 → 80.79 (+5.33)
65.71 → 83.17 (+17.46)
Claude-Opus-4.7
90.67
75.46 → 93.75 (+18.29)
74.83 → 79.29 (+4.46)
63.29 → 81.38 (+18.08)
Gemini-3.1-Pro-Preview
87.25
48.25 → 76.50 (+28.25)
32.25 → 57.42 (+25.17)
64.79 → 76.83 (+12.04)
InternVL3-38B
68.13
48.08 → 74.75 (+26.67)
32.04 → 33.75 (+1.71)
51.04 → 80.12 (+29.09)
Qwen3.6-35B-A3B
72.54
43.58 → 60.58 (+17.00)
20.21 → 38.71 (+18.50)
52.96 → 76.62 (+23.67)
Average
81.58
58.48 → 77.24 (+18.76)
46.96 → 57.99 (+11.03)
59.56 → 79.63 (+20.07)
Appendix
Table 8: Image + Oracle-informed diagnostic control (DCR in percent; gains in percentage points). The supplied reference text is derived from answer fields: Y1 uses the symbolized adapted answer, while Y2 and Y4 use the baseline answer and causal rationale. The Y2 constrained entries use the fixed-800 DCR values from the final harmonized release; all entries remain DCR rather than CSR/eDCR. These are answer-informed diagnostics, not strict edge-only controls. Averages and gains are computed from unrounded scores.
Model
X1 DCR
X2 DCR
X3 DCR
X4 DCR
Avg DCR
GPT-5.5
82.3
92.7
88.2
94.2
89.3
Claude-Opus-4.7
89.7
91.0
89.3
92.7
90.7
Doubao-Seed-2-Pro
91.5
90.0
87.3
90.5
89.8
Gemini-3.1-Pro-Preview
85.7
85.8
88.0
89.5
87.3
Gemini-3.5-Flash
86.0
88.5
91.7
93.0
89.8
GLM-5V-Turbo
71.5
76.5
79.3
82.7
77.5
Appendix
Table 9: Performance on the unconstrained Baseline set. DCR values are percentages. X1 : Causal Discovery, X2 : Forward Prediction, X3 : Retrospective Diagnosis, and X4 : Intervention-Outcome Prediction.
Model
Y1 DCR
Y2 DCR
Y3 DCR
Y4 DCR
eDCR
eCDI (points)
GPT-5.5
77.0
75.5
97.7
65.7
76.9
12.4
Claude-Opus-4.7
75.5
74.8
95.2
63.3
74.5
16.1
Doubao-Seed-2-Pro
72.4
32.4
92.3
53.1
60.0
29.8
Gemini-3.1-Pro-Preview
48.3
32.3
94.7
64.8
58.6
28.6
Gemini-3.5-Flash
48.6
37.6
95.3
58.8
56.2
33.6
GLM-5V-Turbo
58.8
33.5
92.9
63.7
60.4
17.1
Appendix
Table 10: Performance on the constrained main set. The Y1 – Y4 columns report per-constraint DCR, whereas eDCR and eCDI aggregate all four constraints. DCR and eDCR are percentages; eCDI is the baseline-to-effective-score drop in percentage points, computed from unrounded fixed-denominator sample-level scores. Y1 : Entity Symbolization, Y2 : Spatial Grounding, Y3 : Factual Adversarial Constraint, and Y4 : Minimalist Output.
Figure 3: Example of benchmark construction (Image ID: 2317269)
Table 11: Overview of the 14 evaluated MLLMs and the GPT-5.1 automated judge. Open/open-weight parameter sizes are taken from the corresponding official model cards. A numerical value prefixed by ∼ denotes a literature-based estimate of effective knowledge capacity in open-model-equivalent parameters ( Li, 2026 ) , not an officially disclosed parameter count; ∼ alone indicates that the parameter count is not publicly disclosed. Access methods and model sources are reported separately.
Metric
n
Unanimous
Fleiss’ κ
Human–GPT-5.1
L1 perception
280
249/280 (88.93%)
0.7143
227/280 (81.07%)
L2 interaction/mechanism
280
252/280 (90.00%)
0.7848
235/280 (83.93%)
L3 causal outcome
280
255/280 (91.07%)
0.8365
244/280 (87.14%)
Full-chain pass
280
234/280 (83.57%)
0.7415
217/280 (77.50%)
Appendix
Table 12: Independent human validation of the DCR rubric and GPT-5.1 judge. The audit contains 280 responses from 14 models. Unanimous denotes agreement among all three annotators. Human–GPT-5.1 denotes agreement between the human-majority label and the automated judge. Full-chain pass is the binary conjunction L1∧L2∧L3 per annotator; the GPT-5.1 full-chain label is computed from its three binary layer scores.
Setting
n
Human DCR
GPT-5.1 DCR
Human–GPT-5.1 gap
Baseline
56
90.48
72.02
+18.45
Y1
56
63.10
50.00
+13.10
Y2
56
76.19
51.19
+25.00
Y3
56
98.81
93.45
+5.36
Y4
56
63.69
49.40
+14.29
Appendix
Table 13: Setting-level DCR on the 280-response human-audit subset (percent). Human DCR applies the standard cascade after layer-wise majority voting. The gap is human DCR minus GPT-5.1 DCR in percentage points. Values are computed from unrounded scores.
Figure 4: Prompt used for generating causal reasoning questions from human-verified causal-edge specifications and scene-graph annotations. It enforces structured grounding, causal validity, and diversity across four dimensions ( X1 – X4 ).
Figure 5: LLM judge prompts for evaluation. The DCR prompt decomposes causal reasoning into perception (L1), mechanism (L2), and outcome (L3), while the CSR prompt evaluates constraint adherence under contradiction settings ( Y3 ). The prompts are scored separately; under Y3 , their false-premise criteria partially overlap rather than forming a strictly orthogonal decomposition.
CFAR, IHPC, Agency for Science, Technology and Research (A*STAR) Singapore · I2R, Agency for Science, Technology and Research (A*STAR) Singapore · Nanyang Technological University Singapore