cs.CVSep 25, 2026

CCRV-Bench: Constraint-Based Evaluation of Causal Reasoning in Vision-Language Models

Authors: Linyuan Gao, Yuan Wu, Yi Chang

Organizations: School of Artificial Intelligence, Jilin University Changchun 130012, China · Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, Jilin University Changchun 130012, China · International Center of Future Science, Jilin University Changchun 130012, China

Abstract

Vision-language models (VLMs) have demonstrated excellent performance in visual tasks, but their visual causal reasoning capabilities still lack reliable evaluation. Existing evaluations struggle to distinguish whether a model is performing causal reasoning based on visual evidence or relying on statistical correlations for shortcut learning, thereby potentially overestimating their actual capabilities. This paper proposes CCRV-Bench, a constraint-driven visual causal reasoning benchmark for single-image physical scenarios. We construct an orthogonal framework that evaluates four causal task dimensions: causal relation discovery, state prediction, causal diagnosis, and intervention-outcome prediction. We further introduce entity symbolization, spatial grounding, the factual adversarial constraint, and minimalist output constraints to reduce shortcut cues while preserving the physical commonsense required by the task. Experiments across 14 multimodal models show that constraint sensitivity is task- and model-dependent: intervention-outcome prediction has the largest average effective degradation among the four causal tasks, spatial grounding is the most damaging constraint on average, and the factual adversarial constraint improves DCR for all evaluated models. These results show that unconstrained performance does not determine constrained robustness and that a single aggregate score can obscure distinct failures in causal identification, spatial grounding, and constraint-compliant expression. CCRV-Bench provides a standardized framework for diagnosing image-grounded causal reasoning under controlled constraints. The code is available at https://github.com/0815linyuan/CCRVBench

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. What's Missing in Vision-Language Models? Probing Their Struggles with Causal Order Reasoning

    Jun 1, 2025Zhaotian Weng, Haoxuan Li, Xin Eric Wang +2VLM EvaluationVLM Reasoning

  2. From Prompts to Tokens: Internalizing Causal Supervision in Vision-Language Model for Multi-Image Causal Reasoning

    Jun 10, 2026Haoping Yu, Yuanxi Li, Jing MaVisual ReasoningCausal Structure Learning

  3. Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs

    Jun 4, 2026Tianyi Tang, Zhuoyi Lin, Zeyu Feng +4VLM EvaluationVLM Reasoning