We introduce PowerSim, a method to bring physically grounded, differentiable dynamics to PowerFoam's power diagram based 3D representation. PowerSim directly couples a pre-trained PowerFoam scene to the Material Point Method (MPM) by exploiting a natural alignment between the two: the geometric and appearance properties of each primitive correspond closely to the quantities MPM already tracks as an object deforms. Consequently, simulated motion can drive the scene's geometry and appearance directly, without an auxiliary representation in between. Built on this framework, we enable a range of applications on real and synthetic scenes: (1) simulating a static scene under user interaction, (2) recovering spatially varying material fields, (3) compositing primitives from independently captured scenes into a single simulation-ready scene and (4) ray-tracing reflections that update consistently as the object deforms. Our results suggest that PowerSim excels over previous frameworks for physically grounded dynamics, while unlocking unique advantages-such as secondary ray lighting effects on dynamic scenes. Results are best viewed on our project website: https://power-sim.github.io/.
Figures & tables
Figure 2 : PowerSim partitions the reconstructed object volume into explicit cells, while PhysGaussian ( Xie et al., 2023 ) uses Gaussian kernels that tend to concentrate near surfaces and optionally adds interior particles. Under extreme stretching with the same applied force, PAC-NeRF loses texture and develops abrupt cut-like boundaries, while PhysGaussian exhibits pronounced blurring in stretched regions. PowerSim retains more surface detail and sharper boundaries during separation.
Figure 3 : PowerSim overview. After N MPM substeps, we update primitive positions, isotropically scale their radii to match local volume change, and rotate their dipole frames and appearance directions. We then rebuild adjacency for the updated primitives. Insets show two selected primitives before and after deformation.
Figure 4 : Material estimation changes the simulated response. Under the same applied force, random material assignment produces large stem bending, while PhysDreamer produces little bending. PowerSim produces back-and-forth motion with visible stem deformation. The left column shows Young’s modulus; the remaining columns show matching simulation times.
Figure 5 : Scene composition. Our optimization-free primitive selection enables removing the vase from a captured scene (a–b) and selecting a plant from another capture (c, green). We align and insert the selected primitives to create a scene ready for simulation (d).
Figure 6 : Diverse Material Behaviors. Simulation results on different material settings.
Figure 7 : Dynamic reconstruction under deformation and contact. PowerSim preserves the backpack’s texture and Mario’s shape as they fall and contact the ground. PAC-NeRF has blurred textures and distorted silhouettes, while PhysDreamer shows discrepancies in deformation and appearance. Reference frames are shown in the bottom row; time progresses from left to right.
Metrics
Method
backpack
bell
blocks
bus
cream
elephant
grandpa
leather
lion
mario
sofa
turtle
Mean
PSNR ↑
PAC-NeRF
19.37
25.00
23.36
20.72
23.24
22.27
21.63
20.85
22.66
21.01
22.49
22.19
22.06
PhysDreamer
18.93
19.54
19.79
18.82
19.83
17.14
16.91
18.27
18.28
17.56
19.58
23.41
19.00
Ours
23.69
24.39
29.56
26.84
25.82
22.80
22.01
35.33
25.17
20.31
24.72
28.11
25.73
SSIM ↑
PAC-NeRF
0.887
0.956
0.940
0.908
0.893
0.922
0.939
0.932
0.936
0.921
0.926
0.923
0.924
PhysDreamer
0.856
0.940
0.917
0.909
0.876
0.892
0.922
0.945
0.910
0.901
0.894
0.939
0.908
Ours
0.889
0.944
0.953
0.938
0.923
0.908
0.889
0.978
0.930
0.941
0.918
0.953
0.930
Table 1 : Dynamic reconstruction on 12 objects. PowerSim achieves the highest PSNR on 10 objects and SSIM on seven. Bold and underline indicate the best and second-best results, respectively.
Metrics
Method
backpack
bell
blocks
bus
cream
elephant
grandpa
leather
lion
mario
sofa
turtle
Mean
log(E)
PAC-NeRF
3.28
1.08
4.02
3.30
3.22
3.05
2.99
1.20
2.34
3.37
0.20
1.94
2.50
PhysDreamer
0.10
1.07
0.74
0.40
0.28
0.54
0.43
0.33
1.01
1.18
0.68
0.44
0.60
Ours
0.06
0.81
0.16
0.60
0.37
0.36
0.47
1.21
0.20
0.60
0.33
0.09
0.44
ν
PAC-NeRF
0.21
0.23
0.33
0.16
0.12
0.06
0.36
0.26
0.14
0.33
0.30
0.01
0.21
PhysDreamer
0.26
0.01
0.10
0.14
0.14
0.04
0.33
0.11
0.06
0.44
0.05
0.29
0.16
Ours
0.17
0.18
0.16
0.12
0.17
0.27
0.01
0.11
0.16
0.25
0.09
0.21
0.16
Table 2 : Material estimation on 12 objects. MAE in log Young’s modulus and Poisson’s ratio (lower is better). Bold and underline indicate the best and second-best results, respectively.
Figure 8 : Geometry and appearance updates under large deformation. Fixing dipole orientations produces ragged boundaries, while fixing appearance directions introduces color inconsistencies. The full PowerSim update preserves cleaner surface detail; PhysGaussian shows blurring and fragmentation in stretched regions. Insets highlight these differences. Best viewed on supplementary webpage.
Figure 9 : Ray tracing on dynamic scenes. Independently captured objects are composited and simulated together, with mirror reflections that follow their motion and deformation.
Figure 10 : Disocclusion artifacts. PowerSim can select and move the telephone handset, but its motion exposes poorly reconstructed regions, revealing gaps and visual artifacts.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Scene
Figure
Constitutive Model
Bonsai
Fig. 1
Fixed corotated
Bread roll
Fig. 2
Fixed corotated
Carnation
Fig. 4
Fixed corotated
GSO benchmark
Fig. 7
Fixed corotated
Bread twist
Fig. 8
Fixed corotated
Car
Fig. 9
Fixed corotated
Appendix
Table 3 : List of constitutive models used for simulation.
β
garden
bonsai
room
counter
kitchen
average
0.1
0.912 / 0.948
0.837 / 0.881
0.604 / 0.953
0.758 / 0.912
0.855 / 0.947
0.793 / 0.928
0.2
0.916 / 0.948
0.833 / 0.877
0.608 / 0.955
0.758 / 0.912
0.855 / 0.947
0.794 / 0.928
0.4
0.922 / 0.948
0.826 / 0.868
0.615 / 0.953
0.759 / 0.912
0.858 / 0.947
0.796 / 0.926
0.5
0.920 / 0.945
0.822 / 0.863
0.619 / 0.953
0.759 / 0.912
0.859 / 0.947
0.796 / 0.924
0.6
0.920 / 0.944
0.819 / 0.860
0.622 / 0.952
0.759 / 0.912
0.859 / 0.947
0.796 / 0.923
0.8
0.917 / 0.940
0.814 / 0.853
0.625 / 0.948
0.759 / 0.912
0.859 / 0.947
0.795 / 0.920
Appendix
Table 4 : Ablation studies on discounting factor β with mIoU / mAcc per scene.
Method
Opt.-free
Garden
Bonsai
Room
Counter
Kitchen
Average
mIoU ↑ / mAcc ↑
mIoU ↑ / mAcc ↑
mIoU ↑ / mAcc ↑
mIoU ↑ / mAcc ↑
mIoU ↑ / mAcc ↑
mIoU ↑ / mAcc ↑
LabelGS
✗
0.79 / 0.95
0.70 / 0.92
0.64 / 0.93
0.55 / 0.93
0.82 / 0.96
0.70 / 0.94
Gaussian Grouping
✗
0.88 / 0.92
0.75 / 0.82
0.55 / 0.78
0.74 / 0.92
0.83 / 0.92
0.75 / 0.87
SemanticFoam
✗
0.94 / 0.96
0.90 / 0.94
0.63 / 0.95
0.75 / 0.91
0.90 / 0.94
0.82 / 0.94
Ours
✓
0.92 / 0.95
0.82 / 0.86
0.62 / 0.95
0.76 / 0.91
0.86 / 0.95
0.80 / 0.92
Appendix
Table 5 : Per-scene segmentation result on Mip-NeRF 360. Baseline numbers are taken from the SemanticFoam paper. All baselines optimize a per-scene semantic field; ours requires no optimization .
Figure 11 : We provide qualitative results compared with SemanticFoam on Mip-NeRF 360 scenes
To study the ability to infer physical dynamics from videos and extrapolate them forward in time, we assemble a dataset of 2D Material Point Method (MPM) physical simulations covering rich physical phenomena such as deformable objects, fluids, kinetic objects, and emitters. We study code generation and video diffusion approaches on this dataset, identifying their strengths and weaknesses by varying the amount of physically relevant side information. The code generation model, beyond giving a working demonstration of automatic synthesis of MPM simulations, reveals that such an approach struggles with inferring physical parameters from visual input, but relative to video diffusion, produces physically and temporally stable extrapolations forward in time, while the video diffusion model more strongly identifies geometric properties from visual input but produces physically implausible extrapolations.
We introduce a differentiable 3D representation that unifies the ray tracing capabilities of foam-based ray tracing with the efficiency of modern rasterization pipelines. While prior foam representations enable constant-time ray traversal through an explicit volumetric partition of space, their potentially unbounded cells hinder efficient tile-based rasterization. We address this limitation by generalizing Voronoi foams to bounded power diagrams with controllable cell extents, enabling spatially bounded primitives without requiring expensive Delaunay triangulations during training. We further introduce an oriented surface formulation that explicitly models interfaces between interior and exterior regions, and decouple geometry from appearance by embedding differentiable texture directly on these surfaces. Together, these contributions yield a representation that preserves state-of-the-art ray tracing efficiency while achieving rasterization performance competitive with current generation 3DGS, providing a practical path toward unified real-time differentiable rendering.
Shrisudhan Govindarajan, Daniel Rebain, Dor Verbin +3
Simon Fraser University · Google · Google Deepmind +2
Learning physically plausible dynamics from visual observations is essential for interactive world models and embodied agents. However, modeling real-world deformable objects remains challenging because their dynamics often arise from complex, spatially heterogeneous material responses. To address this challenge, we propose PhysReal, a video-driven framework for learning and simulating the underlying physics of real deformable objects. PhysReal integrates a spatially varying hybrid expert-neural constitutive model with a differentiable MPM simulator and 3DGS renderer. Analytical expert models provide interpretable physical priors, while neural constitutive residuals capture material responses beyond predefined formulations. Spatially distributed patches parameterize the constitutive field, enabling a continuous representation of local material variations. To organize the identification of this model from sparse visual observations, we adopt a progressive curriculum that sequentially optimizes global material properties, spatially varying local parameters, and neural constitutive residuals, together with complementary motion and mask supervision. Extensive experiments on diverse deformable-object interactions demonstrate that PhysReal achieves superior performance in dynamic reconstruction and future-state prediction, while showing strong potential for downstream robotic applications.