cs.CVOct 6, 2026

Do Vision Models Learn Physical Constraints or Rendering Shortcuts? A Counterfactual Benchmark for Grounded Physical Consistency

Authors: M. Moein Esfahani, Sepehr Salem, Mohammed Alser, Vince Calhoun

Organizations: Tri-institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS) Georgia State University, Georgia Institute of Technology, and Emory University, Atlanta, GA, USA · Georgia State University, Atlanta, GA, USA

Abstract

Modern image editing models can satisfy a text instruction while breaking the physics of the edited scene. A new object may cast no shadow, a mirror may fail to reflect visible geometry, or an object may float above a surface that should support it. We study physical plausibility diagnosis, detecting whether an edited image violates scene physics, naming the violation type, localizing the affected region, and explaining the failure in language. We introduce a counterfactual benchmark whose controlled synthetic component uses Mitsuba~3 to generate 5,500 images from 500 scene families. Each family contains one clean image and ten matched violations involving shadows, reflection, support, surface response, and occlusion. The renderer pipeline provides category labels, affected-region masks and boxes, scene metadata, and explanation targets. We use LLaVA-1.5-7B, Qwen2.5-VL-7B, and InternVL3.5-8B as diagnostic baselines rather than proposed methods. On a 1,650-image synthetic test set, the adapted baselines reach 64.0--67.8% category macro-F1 on standard held-out scenes. For LLaVA-1.5-7B, category macro-F1 falls from 64.0% on the standard split to 40.8% under intervention shift. This gap shows that high in-distribution accuracy partly reflects cues tied to rendering and counterfactual construction.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing

    Jun 25, 2026Shengbin Guo, Shaokang He, Chaoyue Meng +4Image EditingGenerative Video Models

  2. VideoPhysEdit: Physical Counterfactual Video Editing via Rigid-Body Physical Scene Reconstruction

    Sep 28, 2026Conghan Yue, Yuanjie Chen, Yue Han +4Video EditingScene Understanding

  3. Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing

    Apr 16, 2026Yixuan Ding, Wei Huang, Ruijie Quan +2Image EditingEdit