PhysLDM: Latent Diffusion for High-Fidelity Deformable Simulation
Organizations: S-Lab, Nanyang Technological University, Singapore · Shanghai Artificial Intelligence Laboratory, Shanghai, China
Abstract
Neural simulation of high-fidelity deformable bodies is a foundational challenge in computer graphics and physical AI. Long-horizon prediction for high-resolution 3D volumetric meshes is hard: autoregressive methods are susceptible to error accumulation, while direct multi-frame prediction at native resolution is computationally prohibitive. This motivates a compact spatiotemporal latent representation, which is largely unexplored for mesh-based volumetric physics. Meanwhile, it remains unclear whether deterministic regression or generative diffusion is the more appropriate predictive paradigm. To address these coupled challenges, we introduce PhysLDM, a unified latent-diffusion paradigm for one-shot volumetric deformable simulation. Its core is a holistic spatiotemporal VAE that avoids the "staircase" artifacts of standard temporal compression (as in common video VAEs), achieving ~2.48 mm reconstruction precision on meter-scale scenes at up to 78x token compression. Based on this reliable latent space, we systematically compare regression and diffusion methods. Our experiments uncover a key modeling insight: complex deformable dynamics are often chaotic, and in this regime deterministic regression tends to produce non-physical averages, whereas diffusion better models their distribution. Accordingly, we employ a latent diffusion model that effectively learns from the chaotic data to generate physically plausible trajectories. Trained purely kinematically on an Objaverse-scale dataset, a single PhysLDM generalizes zero-shot to unseen OOD datasets (GSO and Toys4K). Its differentiability further enables efficient solution of inverse problems and higher-order design optimization. To our knowledge, PhysLDM is the first high-fidelity spatiotemporal autoencoder and latent-diffusion paradigm for volumetric deformable dynamics, offering a scalable and robust approach to neural simulation.
Figures & tables
| DCC | Recon | Rigid | Recon | Rigid | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| ST-VAE | mean | mean | med | P90 | rot | trans | mean | rot | trans | |
| Objaverse Held-out | Standard | 0.665 | 5.32 | 3.47 | 12.01 | 0.270 | 0.57 | 6.14 | 0.185 | 0.50 |
| Direct world-coord. † | 0.880 | 6.50 | 4.21 | 15.60 | 0.951 | 3.89 | 7.99 | 0.846 | 4.14 | |
| Holistic (ours) | 0.949 | 2.48 | 1.62 | 8.08 | 0.276 | 0.53 | 4.71 | 0.169 | 0.47 | |
| GSO | Standard | 0.641 | 3.46 | 1.64 | 8.83 | 0.161 | 0.48 | 4.30 | 0.149 | 0.42 |
| Objaverse Held-out | GSO | Toys4K | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Prediction Model | Vol. | SI | KE 1 | Err | Vol. | SI | KE 1 | Err | Vol. | SI | KE 1 | Err |
| Diffusion (ours) | 7.9 | 0.24 | 0.80 | 28.2 | 2.0 | 0.00 | 0.86 | 14.0 | 10.3 | 0.32 | 0.80 | 35.6 |
| Regression | 22.7 | 0.79 | 0.72 | 37.5 | 5.1 | 0.01 | 0.72 | 19.9 | 24.3 | 1.07 | 0.69 | 42.6 |
| MGN-style AR | 55.1 | 3.53 | 0.77 | 80.6 | 47.1 | 0.68 | 0.76 | 78.4 | 57.5 | 4.73 | 0.77 | 79.1 |
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| ST-VAE Small | ST-VAE Base | |
|---|---|---|
| Query tokens | 512 | 512 |
| Window | 32 | 64 |
| Temporal tokens | 4 | 8 |
| Temporal compression | ||
| Hidden width | 512 | 512 |
| Joint latent |
| Scene | Holistic (ours) | Standard | + vel. loss | + spec. loss | Adapt. compr. | Direct world-coord. | |
|---|---|---|---|---|---|---|---|
| 1 | 5,181 | 0.972 | 0.728 | 0.850 | 0.781 | 0.418 | 0.972 |
| 2 | 5,913 | 0.969 | 0.589 | 0.717 | 0.587 | 0.583 | 0.928 |
| 3 | 6,005 | 0.986 | 0.478 | 0.668 | 0.479 | 0.488 | 0.943 |
| 4 | 2,245 | 0.974 | 0.822 | 0.761 | 0.769 | 0.551 | 0.917 |
| 5 | 1,526 | 0.970 | 0.640 | 0.616 | 0.684 | 0.684 | 0.824 |
| Difference between the two runs | of size | of size | of size |
|---|---|---|---|
| (a) mean error, whole trajectory | ( ) | ( ) | ( ) |
| (b) final-frame error | ( ) | ( ) | ( ) |
| Non-chaotic ( scenes) | Chaotic ( scenes) | |||||
|---|---|---|---|---|---|---|
| Vol. | KE 1 | Err | Vol. | KE 1 | Err | |
| Regression | ||||||
| Diffusion, single sample | ||||||
| Diffusion, best of | ||||||
| DCC | Recon | Rigid | Recon | Rigid | ||||||
| ST-VAE | mean | mean | med | P90 | rot | trans | mean | rot | trans | |
| Objaverse Held-out | Standard | 0.665 | 5.32 | 3.47 | 12.01 | 0.270 | 0.57 | 6.14 | 0.185 | 0.50 |
| + velocity loss | 0.842 | 5.34 | 3.15 | 11.60 | 0.831 | 3.33 | 6.25 | 0.638 | 3.88 | |
| + spectral loss | 0.697 | 5.17 | 3.42 | 11.40 | 0.290 | 0.73 | 5.96 | 0.210 | 0.66 | |
| Adaptive compression | 0.511 | 5.86 | 3.65 | 13.72 | 0.440 | 5.09 | 6.59 | 0.314 | 4.30 | |
| Volume | Self-int. | Vel. | Acc. | World err | |||
|---|---|---|---|---|---|---|---|
| Prediction Model | viol. % | tris % | viol. % | viol. % | GT 1 | median (mm) | |
| Objaverse Held-out | ST-VAE reconstruction | 8.5 | 0.12 | 0.8 | 16.9 | 0.91 | 1.97 |
| Diffusion (ours) | 7.9 | 0.24 | 7.0 | 25.8 | 0.80 | 28.2 | |
| Regression | 22.7 | 0.79 | 8.6 | 29.2 | 0.72 | 37.5 | |
| MGN-style AR | 55.1 | 3.53 | 21.4 | 69.0 | 0.77 | 80.6 | |
| GSO | ST-VAE reconstruction | 2.1 | 0.00 | 0.5 | 11.8 | 0.92 | 1.21 |
| no ST-VAE | with ST-VAE | |
|---|---|---|
| LDM on raw tokens (GiB) | ST-VAE Base step (GiB) | |
| 500 | 28.9 | 33.5 |
| 1,000 | 55.4 | 37.4 |
| 2,000 | 108.3 | 45.8 |
| 2,500 | 134.5 | 50.7 |
| 2,600 | OOM | 50.7 |
| Decode (ms) | bf16 fps | fp32 fps | Peak | ||||||
|---|---|---|---|---|---|---|---|---|---|
| bf16 | fp32 | Regression | LDM-20 | LDM-50 | Regression | LDM-20 | LDM-50 | (GiB) | |
| 980 | 318 | 288 | 195 | 82 | 44 | 192 | 30 | 13 | 4.8 |
| 2,493 | 327 | 425 | 189 | 79 | 42 | 136 | 28 | 13 | 4.8 |
| 5,007 | 328 | 597 | 188 | 81 | 43 | 100 | 26 | 12 | 4.8 |
| 7,534 | 332 | 778 | 186 | 81 | 43 | 78 | 25 | 12 | 5.2 |
| 9,851 | 384 | 960 | 162 | 76 | 42 | 64 | 23 | 12 | 5.4 |
| Figure | Object | Dataset | (Pa) | (m/s) | |
|---|---|---|---|---|---|
| Fig. 3 | phone | Objaverse (held-out) | 6,106 | , | |
| Fig. 3 | boot | GSO | 6,569 | (released from rest) | |
| Fig. 3 | chess piece | Toys4K | 6,322 | (released from rest) | |
| Fig. 3 | toy | GSO | 9,403 | (released from rest) | |
| Fig. 9 | cake | Toys4K | 6,640 | , | |
| Fig. 9 | shape-sorter toy | GSO | 6,426 | (released from rest) |
| Figure | Object | Dataset | (Pa) | (m/s) | |
|---|---|---|---|---|---|
| Fig. 4 | punch-drop toy | GSO | 7,767 | , | |
| Fig. 4 | cap | Objaverse (held-out) | 2,435 | (released from rest) | |
| Fig. 4 | bowl | Objaverse (held-out) | 5,877 | (released from rest) | |
| Fig. 4 | grapes | Toys4K | 8,767 | (released from rest) | |
| Fig. 11 | bucket | Objaverse (held-out) | 9,853 | , | |
| Fig. 11 | sofa | Objaverse (held-out) | 4,228 | , |
| Simulator | Differentiation Support | 1st-Order Task (Material Estimation) | 2nd-Order Task (Throw Design) |
|---|---|---|---|
| Ours | Autograd (higher order) | Success (40 seconds) | Success (7 minutes) Efficient 2nd-order grads. |
| DiffIPC | Analytic adjoint (1st order only) | 26 minutes Unstable under defaults. Needs manual tuning. | Infeasible (Multi-day) slower via finite diff. |
| Newton/Warp | Recorded tape | Fails Gradients violently explode. | Infeasible No usable 1st-order gradients. |
| tokens | Concurrent (DiT-L) | Concurrent (DiT-B) | |
|---|---|---|---|
| ( ) | GiB / ms | GiB / ms | |
| 500 | 32k | 64.4 / 669 | 24.0 / 212 |
| 1,000 | 64k | 122.1 / 1317 | 46.0 / 412 |
| 1,200 | 77k | OOM | 54.7 / 495 |
| 2,000 | 128k | OOM | 89.9 / 879 |
| 2,500 | 160k | OOM | 111.8 / 1147 |
| Model | Stage | Training time |
|---|---|---|
| Concurrent (DiT-B) | (single) | 74.6 h |
| Ours | spatial VAE | 1.4 h |
| ST-VAE | 40.0 h | |
| LDM | 19.6 h | |
| total | 61.1 h |
| Data | Output | Scenes | Vol. (%) | Err (mm) | |
|---|---|---|---|---|---|
| Stable Neo-Hookean | ST-VAE reconstruction | ||||
| Stable Neo-Hookean | LDM prediction | ||||
| ARAP | ST-VAE reconstruction | ||||
| ARAP | LDM prediction | ||||
| Real captured | ST-VAE reconstruction | – | – | ||
| Real captured | LDM prediction | – | – |
| Vol. (%) | Err (mm) | ||||||
|---|---|---|---|---|---|---|---|
| Data | Scenes | before | after | before | after | before | after |
| Stable Neo-Hookean | |||||||
| ARAP | |||||||
| Real captured | – | – | – | – | |||