Do Vision Models Learn Physical Constraints or Rendering Shortcuts? A Counterfactual Benchmark for Grounded Physical Consistency
Authors: M. Moein Esfahani, Sepehr Salem, Mohammed Alser, Vince Calhoun
Organizations: Tri-institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS) Georgia State University, Georgia Institute of Technology, and Emory University, Atlanta, GA, USA · Georgia State University, Atlanta, GA, USA
Modern image editing models can satisfy a text instruction while breaking the physics of the edited scene. A new object may cast no shadow, a mirror may fail to reflect visible geometry, or an object may float above a surface that should support it. We study physical plausibility diagnosis, detecting whether an edited image violates scene physics, naming the violation type, localizing the affected region, and explaining the failure in language. We introduce a counterfactual benchmark whose controlled synthetic component uses Mitsuba~3 to generate 5,500 images from 500 scene families. Each family contains one clean image and ten matched violations involving shadows, reflection, support, surface response, and occlusion. The renderer pipeline provides category labels, affected-region masks and boxes, scene metadata, and explanation targets. We use LLaVA-1.5-7B, Qwen2.5-VL-7B, and InternVL3.5-8B as diagnostic baselines rather than proposed methods. On a 1,650-image synthetic test set, the adapted baselines reach 64.0--67.8% category macro-F1 on standard held-out scenes. For LLaVA-1.5-7B, category macro-F1 falls from 64.0% on the standard split to 40.8% under intervention shift. This gap shows that high in-distribution accuracy partly reflects cues tied to rendering and counterfactual construction.
Figures & tables
Figure 1: Overview of the benchmark and diagnostic baselines. (A) Controlled rendering creates clean scenes, physical counterfactuals, and category, region, and explanation supervision. (B) LLaVA-1.5-7B, Qwen2.5-VL-7B, and InternVL3.5-8B use the same input and structured output. These models are benchmark probes, not proposed architectures. The evaluation splits are defined in Table 1 .
Partition
Scenes
Images
Role
Training
300
3,300
adaptation
Validation
50
550
model selection
Standard held-out
50
550
unseen scene seeds
Appearance shift
50
550
unseen visual factors
Intervention shift
50
550
unseen construction rules
Total
500
5,500
Table 1: Scene-disjoint synthetic benchmark partitions. Every scene family contains one image for each of the 11 labels.
Method
Bin. Acc
Bal. Acc
Bin. mF1
Clean R
Viol. R
Cat. Acc
Cat. mF1
Majority class
90.9
50.0
47.6
0.0
100.0
9.1
1.5
CLIP zero-shot
13.2
27.7
13.0
45.3
10.0
6.2
1.9
LLaVA-1.5-7B zero-shot
88.4
49.5
48.4
2.0
97.0
1.2
1.2
LLaVA-1.5-7B + LoRA
98.4
92.7
94.8
85.8
99.6
63.7
64.0
Qwen2.5-VL-7B + LoRA
98.6
92.5
95.3
85.1
99.9
68.1
67.8
InternVL3.5-8B + LoRA
97.9
91.2
93.3
83.0
99.4
66.3
66.0
Table 2: Diagnostic baselines on the 550-image standard held-out split. All scores are percentages; mF1 denotes macro-F1. Adapted-model entries are means over three seeds.
Setting
Bin. mF1
Cat. mF1
Box IoU (%)
Invalid (%)
Expl./QA correct (%)
Standard held-out
94.8
64.0
60.5
2.7
77.4
Appearance shift
86.7
52.4
50.6
5.1
68.3
Intervention shift
77.9
40.8
41.7
7.6
58.9
PICABench zero-shot
n/a
n/a
22.7
14.2
38.6
PICABench adapted (+PICA-20K)
n/a
n/a
34.8
8.6
49.1
Human reference (standard)
96.0
82.0
74.0
n/a
91.0
Table 3: LLaVA-1.5-7B + LoRA across evaluation settings, with human performance as a reference. Scores are percentages and average three seeds. n/a denotes a metric incompatible with the native PICABench task.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Additional clean and violated examples for all ten fine-grained categories. Columns vary object geometry, camera view, lighting, and scene appearance. The panel is also used for visual auditing: cues such as hard boundaries, halos, or unrelated color changes are recorded rather than hidden.
Fine label
Operational distinction
Main shortcut risk
Wrong shadow direction
Shadow direction changes while object shading and scene lighting are fixed.
Transferred-shadow boundary
Missing cast shadow
The object and contact cue remain, but its cast shadow is removed.
Local smoothing or removal trace
Floating object
A visible gap separates the object from its support; other lighting remains valid.
Gap or silhouette alone
Shadow–illumination inconsistency
Object shading comes from one light setting and the cast shadow from another.
Close to wrong shadow direction
Missing reflection
A visible object that should appear in a verified reflective region is absent.
Reflection-region texture change
Reflection angle mismatch
Reflected content is displaced from its mirror-consistent position.
Composite edge or displacement
Appendix
Table 4: Operational category definitions. The close shadow and geometry pairs are intentional stress tests. Human agreement and family-level metrics will be reported so that these sub-types are not treated as independent facts.
Category
Recall
F1
Floating object
0.993
0.997
Incorrect specular highlight
0.993
0.997
Missing contact shadow
0.980
0.990
Wrong reflected object
0.993
0.987
Clean
0.867
0.919
Missing cast shadow
1.000
0.804
Appendix
Table 5: Per-category recall and F1 on the 1,650-image development pilot.
Figure 3: Development-pilot confusion matrix. Shadow–illumination inconsistency is absorbed by wrong shadow direction, while impossible occlusion is absorbed by floating object. These close pairs motivate the operational definitions in Table 4 .