Do Vision Models Learn Physical Constraints or Rendering Shortcuts? A Counterfactual Benchmark for Grounded Physical Consistency
Authors: M. Moein Esfahani, Sepehr Salem, Mohammed Alser, Vince Calhoun
Organizations: Tri-institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS) Georgia State University, Georgia Institute of Technology, and Emory University, Atlanta, GA, USA · Georgia State University, Atlanta, GA, USA
Modern image editing models can satisfy a text instruction while breaking the physics of the edited scene. A new object may cast no shadow, a mirror may fail to reflect visible geometry, or an object may float above a surface that should support it. We study physical plausibility diagnosis, detecting whether an edited image violates scene physics, naming the violation type, localizing the affected region, and explaining the failure in language. We introduce a counterfactual benchmark whose controlled synthetic component uses Mitsuba~3 to generate 5,500 images from 500 scene families. Each family contains one clean image and ten matched violations involving shadows, reflection, support, surface response, and occlusion. The renderer pipeline provides category labels, affected-region masks and boxes, scene metadata, and explanation targets. We use LLaVA-1.5-7B, Qwen2.5-VL-7B, and InternVL3.5-8B as diagnostic baselines rather than proposed methods. On a 1,650-image synthetic test set, the adapted baselines reach 64.0--67.8% category macro-F1 on standard held-out scenes. For LLaVA-1.5-7B, category macro-F1 falls from 64.0% on the standard split to 40.8% under intervention shift. This gap shows that high in-distribution accuracy partly reflects cues tied to rendering and counterfactual construction.
Figures & tables
Figure 1: Overview of the benchmark and diagnostic baselines. (A) Controlled rendering creates clean scenes, physical counterfactuals, and category, region, and explanation supervision. (B) LLaVA-1.5-7B, Qwen2.5-VL-7B, and InternVL3.5-8B use the same input and structured output. These models are benchmark probes, not proposed architectures. The evaluation splits are defined in Table 1 .
Partition
Scenes
Images
Role
Training
300
3,300
adaptation
Validation
50
550
model selection
Standard held-out
50
550
unseen scene seeds
Appearance shift
50
550
unseen visual factors
Intervention shift
50
550
unseen construction rules
Total
500
5,500
Table 1: Scene-disjoint synthetic benchmark partitions. Every scene family contains one image for each of the 11 labels.
Method
Bin. Acc
Bal. Acc
Bin. mF1
Clean R
Viol. R
Cat. Acc
Cat. mF1
Majority class
90.9
50.0
47.6
0.0
100.0
9.1
1.5
CLIP zero-shot
13.2
27.7
13.0
45.3
10.0
6.2
1.9
LLaVA-1.5-7B zero-shot
88.4
49.5
48.4
2.0
97.0
1.2
1.2
LLaVA-1.5-7B + LoRA
98.4
92.7
94.8
85.8
99.6
63.7
64.0
Qwen2.5-VL-7B + LoRA
98.6
92.5
95.3
85.1
99.9
68.1
67.8
InternVL3.5-8B + LoRA
97.9
91.2
93.3
83.0
99.4
66.3
66.0
Table 2: Diagnostic baselines on the 550-image standard held-out split. All scores are percentages; mF1 denotes macro-F1. Adapted-model entries are means over three seeds.
Setting
Bin. mF1
Cat. mF1
Box IoU (%)
Invalid (%)
Expl./QA correct (%)
Standard held-out
94.8
64.0
60.5
2.7
77.4
Appearance shift
86.7
52.4
50.6
5.1
68.3
Intervention shift
77.9
40.8
41.7
7.6
58.9
PICABench zero-shot
n/a
n/a
22.7
14.2
38.6
PICABench adapted (+PICA-20K)
n/a
n/a
34.8
8.6
49.1
Human reference (standard)
96.0
82.0
74.0
n/a
91.0
Table 3: LLaVA-1.5-7B + LoRA across evaluation settings, with human performance as a reference. Scores are percentages and average three seeds. n/a denotes a metric incompatible with the native PICABench task.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Additional clean and violated examples for all ten fine-grained categories. Columns vary object geometry, camera view, lighting, and scene appearance. The panel is also used for visual auditing: cues such as hard boundaries, halos, or unrelated color changes are recorded rather than hidden.
Fine label
Operational distinction
Main shortcut risk
Wrong shadow direction
Shadow direction changes while object shading and scene lighting are fixed.
Transferred-shadow boundary
Missing cast shadow
The object and contact cue remain, but its cast shadow is removed.
Local smoothing or removal trace
Floating object
A visible gap separates the object from its support; other lighting remains valid.
Gap or silhouette alone
Shadow–illumination inconsistency
Object shading comes from one light setting and the cast shadow from another.
Close to wrong shadow direction
Missing reflection
A visible object that should appear in a verified reflective region is absent.
Reflection-region texture change
Reflection angle mismatch
Reflected content is displaced from its mirror-consistent position.
Composite edge or displacement
Appendix
Table 4: Operational category definitions. The close shadow and geometry pairs are intentional stress tests. Human agreement and family-level metrics will be reported so that these sub-types are not treated as independent facts.
Category
Recall
F1
Floating object
0.993
0.997
Incorrect specular highlight
0.993
0.997
Missing contact shadow
0.980
0.990
Wrong reflected object
0.993
0.987
Clean
0.867
0.919
Missing cast shadow
1.000
0.804
Appendix
Table 5: Per-category recall and F1 on the 1,650-image development pilot.
Figure 3: Development-pilot confusion matrix. Shadow–illumination inconsistency is absorbed by wrong shadow direction, while impossible occlusion is absorbed by floating object. These close pairs motivate the operational definitions in Table 4 .
While instruction-based image editing, enabled by multi-modal generative models, has advanced significantly, existing benchmarks lack a comprehensive evaluation of physics-based reasoning, a critical capability for handling real-world scenarios. To address this, we introduce PhyEditBench, a benchmark designed to assess the physical understanding of editing models. Guided by a hierarchical taxonomy, we establish 4 primary classes and 12 subclasses. It comprises 238 high-quality, high-resolution, real-world instances meticulously extracted from videos to capture authentic physical dynamics, alongside 35 synthetic Anti-Physics instances. Our empirical analysis of current SOTA editing methods exposes substantial limitations in their physics-based reasoning. We further propose a training-free baseline named PhyWorld that uses test-time scaling and a latent reduction strategy. PhyWorld outperforms comparable models and suggests that the video generation process can effectively serve as a reasoning mechanism for image editing. The project page is available at https://github.com/Previsior/PhyEditBench.
Video editing has advanced substantially in recent years, with methods increasingly accounting for the visual consequences of edits, such as changes to shadows and occlusions. However, the physical consequences of edits, including changes to subsequent motion and interactions, remain less explored. We formulate this problem as physical counterfactual video editing (PCVE), which aims to generate a counterfactual video depicting the resulting motion and interactions given a source video, a physical edit, and its execution frame. PCVE is challenging because it requires understanding scene physics and inferring the downstream motion and interactions induced by a physical intervention, while paired factual and counterfactual data and dedicated evaluation metrics are lacking. We introduce VideoPhysEdit, a new training-free pipeline for PCVE in rigid-body scenes. It makes physical reasoning explicit through a novel physical scene reconstruction method that recovers a scene reproducing the observed motion and interactions under simulation, enabling the pipeline to apply physical edits as interventions and use the resulting trajectories to guide counterfactual video generation. We further construct PCVE-RigidBench, a synthetic benchmark with paired source and counterfactual target videos and physical ground truth, and introduce the Physical Edit Score. VideoPhysEdit achieves substantially higher physical edit accuracy than open-source methods and commercial models while maintaining competitive visual fidelity. Its Physical Edit Score is 0.376, the only positive score among the compared methods. Qualitative comparisons on real videos further show that VideoPhysEdit applies to real-world scenes and better depicts the downstream motion and interactions induced by the edits than the compared methods. Code: https://github.com/Hammour-steak/VideoPhysEdit
Conghan Yue, Yuanjie Chen, Yue Han +4
Institute of Trustworthy Embodied AI, Fudan University
Diffusion-based image editing has achieved strong visual fidelity under natural language instructions, yet most existing systems still operate at the level of surface instruction following, without reasoning about the implicit contextual constraints embedded in real user requests. This often leads to visually plausible but logically inconsistent edits. In this work, we introduce RE-Edit, a benchmark for REasoning-aware image Editing that evaluates image editing systems across five complementary reasoning dimensions: physical, environmental, cultural, causal, and referential. RE-Edit comprises 1,000 carefully curated samples, each designed such that visual plausibility alone is insufficient and correct editing requires satisfying implicit logical constraints. To support fine-grained analysis, we establish dimension-aligned evaluation criteria and conduct a comprehensive study of ten open-source and two commercial image editing models. Our results show that even advanced systems frequently struggle with implicit multi-dimensional reasoning despite producing high-quality visuals. We further present a lightweight reasoning-guided post-edit baseline as an initial exploration, illustrating how inserting explicit reasoning can help mitigate such failures in a model-agnostic manner.