VideoPhysEdit: Physical Counterfactual Video Editing via Rigid-Body Physical Scene Reconstruction
Organizations: Institute of Trustworthy Embodied AI, Fudan University
Abstract
Video editing has advanced substantially in recent years, with methods increasingly accounting for the visual consequences of edits, such as changes to shadows and occlusions. However, the physical consequences of edits, including changes to subsequent motion and interactions, remain less explored. We formulate this problem as physical counterfactual video editing (PCVE), which aims to generate a counterfactual video depicting the resulting motion and interactions given a source video, a physical edit, and its execution frame. PCVE is challenging because it requires understanding scene physics and inferring the downstream motion and interactions induced by a physical intervention, while paired factual and counterfactual data and dedicated evaluation metrics are lacking. We introduce VideoPhysEdit, a new training-free pipeline for PCVE in rigid-body scenes. It makes physical reasoning explicit through a novel physical scene reconstruction method that recovers a scene reproducing the observed motion and interactions under simulation, enabling the pipeline to apply physical edits as interventions and use the resulting trajectories to guide counterfactual video generation. We further construct PCVE-RigidBench, a synthetic benchmark with paired source and counterfactual target videos and physical ground truth, and introduce the Physical Edit Score. VideoPhysEdit achieves substantially higher physical edit accuracy than open-source methods and commercial models while maintaining competitive visual fidelity. Its Physical Edit Score is 0.376, the only positive score among the compared methods. Qualitative comparisons on real videos further show that VideoPhysEdit applies to real-world scenes and better depicts the downstream motion and interactions induced by the edits than the compared methods. Code: https://github.com/Hammour-steak/VideoPhysEdit
Figures & tables
| Method | Physical Edit Accuracy | Visual Fidelity | ||||||
|---|---|---|---|---|---|---|---|---|
| PES | TE | Mask IoU | PSNR | SSIM | LPIPS | CLIP | FVD | |
| VACE | 146.26 | 0.273 | 14.06 | 0.728 | 0.447 | 0.795 | 1184.37 | |
| Ditto | 149.94 | 0.264 | 22.07 | 0.812 | 0.206 | 0.850 | 551.14 | |
| MiniMax H3 | 152.00 | 0.250 | 24.87 | 0.870 | 0.127 | 0.906 | 246.45 | |
| Seedance 2.5 | 144.99 | 0.231 | 28.85 | 0.928 | 0.080 | 0.932 | 246.67 | |
| No edit | 0.000 | 143.13 | 0.289 | 31.23 | 0.974 | 0.036 | 0.957 | 249.68 |
| Method | Physical Edit Accuracy | Visual Fidelity | ||||||
|---|---|---|---|---|---|---|---|---|
| PES | TE | Mask IoU | PSNR | SSIM | LPIPS | CLIP | FVD | |
| VACE | 140.72 | 0.307 | 12.34 | 0.651 | 0.503 | 0.762 | 1782.50 | |
| Ditto | 153.98 | 0.255 | 21.13 | 0.805 | 0.246 | 0.790 | 892.24 | |
| MiniMax H3 | 0.383 | 96.24 | 0.361 | 25.23 | 0.889 | 0.104 | 0.914 | 256.52 |
| Seedance 2.5 | 0.290 | 106.76 | 0.245 | 27.86 | 0.892 | 0.085 | 0.927 | 272.93 |
| VOID | 0.394 | 84.07 | 0.314 | 29.22 | 0.914 | 0.168 | 0.862 | 262.43 |
| Output | Metric | Result |
| Stage 3 4 | Mask IoU | |
| Stage 5 init. opt. | Stage 5 Mask IoU | |
| Stage 6 PES | ||
| Stage 6 7 | PES | |
| TE | ||
| Mask IoU |
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
| Edit | Tasks |
|---|---|
| Mass | 32 |
| Friction | 21 |
| Restitution | 13 |
| Initial velocity | 26 |
| Add | 7 |
| Delete | 30 |
| Method | Resolution | Output frames | Evaluated frames | Tasks | Settings |
|---|---|---|---|---|---|
| VACE | 96 | 96 | 129 | 30 steps, CFG 5.0 | |
| Ditto | 73 | 73 | 129 | VACE-14B with Ditto LoRA | |
| MiniMax H3 | 768p | 107 | 96 | 129 | video editing API |
| Seedance 2.5 | 720p | 89 | 89 | 129 | video editing API |
| VOID | 96 | 96 | 30 | removal only | |
| VideoPhysEdit | 96 | 96 | 129 | 16 steps, CFG 1.0 |
| Evaluation set | Tasks | Physical Edit Accuracy | Visual Fidelity | ||||||
|---|---|---|---|---|---|---|---|---|---|
| PES | TE | Mask IoU | PSNR | SSIM | LPIPS | CLIP | FVD | ||
| Generated videos | 116 | 0.418 | 63.68 | 0.421 | 27.47 | 0.921 | 0.107 | 0.928 | 198.79 |
| No edit | 129 | 0.000 | 143.13 | 0.289 | 31.23 | 0.974 | 0.036 | 0.957 | 249.68 |
| Complete benchmark | 129 | 0.376 | 66.70 | 0.421 | 27.51 | 0.925 | 0.104 | 0.929 | 182.46 |
| Execution timing | Tasks | PES | TE | Mask IoU |
|---|---|---|---|---|
| First frame | 117 | 0.356 | 68.15 | 0.403 |
| Partway | 12 | 0.571 | 52.57 | 0.602 |
| Method | Add | Delete | Set |
|---|---|---|---|
| VACE | |||
| Ditto | |||
| MiniMax H3 | 0.383 | ||
| Seedance 2.5 | 0.290 | ||
| No edit | 0.000 | 0.000 | 0.000 |
| VideoPhysEdit | 0.032 | 0.633 | 0.318 |
| Method | Directly edited | Other affected | All measurable | |||
|---|---|---|---|---|---|---|
| PES | TE | PES | TE | PES | TE | |
| VACE | 179.32 | 122.79 | 146.26 | |||
| Ditto | 165.86 | 139.13 | 149.94 | |||
| MiniMax H3 | 0.009 | 161.44 | 140.13 | 152.00 | ||
| Seedance 2.5 | 166.37 | 126.88 | 144.99 | |||
| VideoPhysEdit | 0.451 | 70.82 | 0.276 | 72.15 | 0.376 | 66.70 |
| Method | Friction | Initial velocity | Mass | Presence | Restitution |
|---|---|---|---|---|---|
| VACE | |||||
| Ditto | |||||
| MiniMax H3 | 0.212 | ||||
| Seedance 2.5 | 0.208 | ||||
| No edit | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 |
| VideoPhysEdit | 0.315 | 0.435 | 0.307 | 0.520 | 0.116 |
| Stage | Mean | Median | 25th percentile | 75th percentile |
|---|---|---|---|---|
| Stage 3 canonical and anchor scenes | 0.883 | 0.899 | 0.832 | 0.945 |
| Stage 4 motion prior | 0.887 | 0.908 | 0.861 | 0.944 |
| Stage 5 physical rollout | 0.678 | 0.733 | 0.602 | 0.799 |
| Stage | PES | TE | Mask IoU |
|---|---|---|---|
| Stage 6 projected trajectories | 0.398 | 63.39 | 0.391 |
| Stage 7 generated video | 0.412 | 64.76 | 0.412 |
| Stage | Observation | Method |
|---|---|---|
| 2 | Unreliable or missing masks | Weight motion estimates by observation confidence and identify stable intervals, transition episodes, and unresolved observations |
| 3 | Fragmented support planes | Merge compatible planes, refit their combined 3D points, and refine finite boundaries using image outlines |
| 3 | Sparse 3D correspondences | Use 2D correspondences and dense object points to supplement sparse 3D correspondences when fitting pose and scale |
| 3 | Depth scale across frames | Align each motion anchor frame to the canonical scene using the static background and keep object scale fixed |
| 3 | Multiple support assignments | Check contact, nonpenetration, and support plane boundaries; reconcile support relations across motion anchor frames |
| 4 | Missing motion observations | Fit motion models to stable intervals and connect them with transition curves constrained by boundary states |
| Variant | Stage 5 Mask IoU | Stage 6 PES | Stage 6 TE | Stage 6 Mask IoU |
|---|---|---|---|---|
| Calibrated initialization | 0.375 | 0.269 | 76.70 | 0.324 |
| Search and refinement | 0.678 | 0.403 | 62.37 | 0.392 |
| Stage | Average time (s) | Peak GPU memory |
|---|---|---|
| 1 Object identification and tracking | 32.7 | 4.1 GiB |
| 2 Motion observation analysis | 27.6 | 11.9 GiB |
| 3 Canonical and anchor scene reconstruction | 117.5 | 12.0 GiB |
| 4 Motion prior reconstruction | 70.1 | CPU only |
| 5 Physical inversion | 246.7 | CPU only |
| 6 Physical intervention | 22.6 | depends on operation |
| Method | Heavy block | Grippy car | Heavy red ball | Strong push |
|---|---|---|---|---|
| Seedance 2.5 | ||||
| MiniMax H3 | ||||
| VACE | ||||
| Ditto |