Organizations: NLPR, MAIS, Institute of Automation, Chinese Academy of Sciences · School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Independent Researcher
Deformable-object manipulation is essential for robotic tasks such as folding laundry and handling food, where robots must control shape changes as well as object motion. Predictive soft-body simulation supports these tasks by anticipating deformation under external interactions. However, spatial neighborhoods can misrepresent deformation dependencies, introducing local errors that accumulate over successive predictions. Models fitted to individual scenes must also accommodate changes in object geometry and manipulation conditions. In this work, we propose AIM, an Adaptive Interaction Modeling framework that treats real-to-sim soft-body simulation as a local-global interaction modeling problem. AIM uses motion history and geometry to adapt particle relations over current spatial neighbors and retained connections, while geometry-conditioned global communication coordinates object-wide responses. A unified kinematic control-point interface represents different manipulation configurations, and multi-step autoregressive supervision trains the model on its own predicted trajectories. Experiments on PhysTwin and PGND demonstrate improved motion accuracy and visual fidelity, with a 20.0% reduction in future-prediction tracking error relative to PhysTwin and a 22.8% reduction in mean long-horizon particle error across six object categories relative to PGND. The framework further supports transfer across actions, object instances, and scenes, including zero-shot transfer from robot interactions to human manipulation without target-domain dynamics fitting.
Figures & tables
Figure 1: Challenges in real-to-sim deformable dynamics.
Figure 2: Overview of AIM . Adaptive particle relations and geometry-conditioned global communication jointly predict object dynamics under kinematic controls.
Reconstruction & Resimulation
Future Prediction
Method
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
PhysTwin
0.005
0.010
84.4
28.636
0.949
0.030
0.012
0.025
68.2
25.321
0.942
0.056
SoMA
0.018
0.058
67.0
25.007
0.917
0.100
0.041
0.148
39.5
21.776
0.865
0.136
PGND
0.009
0.013
82.1
28.228
0.949
0.033
0.024
0.044
60.3
24.426
0.938
0.069
AIM
0.006
0.007
87.0
29.473
0.952
0.025
0.014
0.020
69.8
25.916
0.943
0.055
Table 1: Quantitative comparison on reconstruction & resimulation and future prediction on the PhysTwin benchmark. Best results are shown in bold , and second-best results are underlined .
Figure 3: Qualitative Results on Reconstruction & Resimulation and Future Prediction.
Figure 4: Qualitative Results on Held-out Interaction Prediction.
Figure 5: Long-Horizon Rollout Performance on Held-out Interaction Prediction.
Cross-Action Transfer
Cross-Object Transfer
Method
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
PhysTwin
0.035
0.053
53.1
21.679
0.900
0.110
–
–
–
–
–
–
PGND
0.226
0.274
58.4
23.763
0.902
0.089
0.026
0.038
72.2
23.023
0.934
0.086
AIM
0.020
0.045
70.6
24.430
0.904
0.073
0.018
0.023
79.8
24.506
0.942
0.073
Table 2: Quantitative comparison on cross-action transfer and cross-object transfer on the PhysTwin benchmark. Best results are shown in bold . “–” indicates that transferring PhysTwin’s object-specific spring-mass model requires additional adaptation beyond our evaluation protocol.
Held-Out Scene Generalization
Held-Out Action Generalization: Lift → Fold
Method
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
PGND
0.023
–
84.3
24.118
0.923
0.090
0.032
–
69.1
22.037
0.900
0.109
AIM
0.013
–
87.1
24.869
0.929
0.074
0.019
–
71.3
22.982
0.896
0.108
Table 3: Generalization to held-out cloth scenes and unseen action types . Best results are shown in bold . “–” indicates unavailable ground-truth 3D tracking annotations.
Method
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
PGND
0.019
0.033
60.28
24.835
0.942
0.070
AIM
0.016
0.023
65.21
25.438
0.941
0.069
Table 4: Zero-shot transfer from PGND to PhysTwin. Best results are shown in bold .
Figure 6: Ablation study on double_stretch_zebra
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Additional qualitative comparisons on PhysTwin for reconstruction and resimulation (left) and future prediction (right). Red boxes highlight local deformation differences.
Object
Method
MDE ↓
CD ↓
EMD ↓
IoU ↑
Boundary F ↑
PSNR ↑
SSIM ↑
LPIPS ↓
PGND
45.80
43.07
22.05
65.40
44.22
23.86
0.916
0.108
Cloth
AIM
44.15
38.67
18.45
66.13
43.95
23.92
0.919
0.106
PGND
47.52
47.47
26.35
25.43
46.68
23.34
0.975
0.042
Rope
AIM
22.49
20.63
10.87
32.24
57.50
23.39
0.977
0.034
PGND
45.07
34.36
18.03
58.90
47.20
23.40
0.932
0.069
Plush
AIM
38.71
29.89
14.70
58.62
50.30
23.44
0.935
0.067
Appendix
Table 5: Quantitative comparison with PGND on unseen interaction prediction across six deformable object categories. Best results are shown in bold .
Figure 8: Qualitative Results on Cross-Action Transfer on the Same Object.
Method
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Cloth 3: Double lift → Single clift
PhysTwin
0.010
0.029
88.6
24.310
0.877
0.126
PGND
0.011
0.023
94.9
27.932
0.882
0.100
AIM
0.011
0.022
95.3
28.543
0.886
0.047
Sloth: Double lift → Double stretch
PhysTwin
0.051
0.063
35.4
19.662
0.899
0.132
Appendix
Table 6: Per-scene results on cross-action transfer. Each model is transferred from the source interaction to the target interaction without target-specific training. Best results within each transfer pair are shown in bold .
Figure 9: Qualitative Results on Cross-Object Transfer under Matched Action Types.
Method
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Single lift: Cloth 1 → Cloth 4
PGND
0.040
0.053
62.4
20.332
0.922
0.107
AIM
0.023
0.029
76.9
23.066
0.939
0.081
Double lift: Cloth 3 → Cloth 1
PGND
0.012
0.021
81.9
25.715
0.945
0.065
AIM
0.012
0.016
82.6
25.945
0.944
0.064
Appendix
Table 7: Per-scene results on cross-object transfer under matched action types. Best displayed results within each transfer pair are shown in bold .
Method
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Rope: robot → human manipulation
PGND
0.014
0.029
39.70
26.132
0.976
0.035
AIM
0.013
0.017
54.00
27.738
0.977
0.030
Cloth: robot → human manipulation
PGND
0.021
0.032
76.60
25.130
0.926
0.088
AIM
0.017
0.022
78.50
25.561
0.927
0.085
Appendix
Table 8: Per-scene results on zero-shot transfer from PGND to PhysTwin. Best results within each scene are shown in bold .
Simulating deformable objects is essential for a wide range of robotic manipulation applications, yet accurately predicting their dynamics remains challenging. We propose Physics-Guided Residual Dynamics (PGRD), a hybrid simulation framework that combines the advantages of physics-based and learning-based approaches. Specifically, PGRD combines an optimizable spring-mass simulator as a backbone with a learned neural network that predicts residual corrections to the physics-based predictions. We adopt a velocity-based formulation to ensure stable simulation and a sliding-window transformer architecture to capture temporal dependencies. We show that PGRD produces more accurate results than both purely physics-based and learning-based methods on a set of diverse real-world deformable objects. We further demonstrate the utility of PGRD in two applications: manipulation planning via Model Predictive Control, including a language-conditioned setting with a generated goal image; and interactive simulation via action-conditioned video prediction by 3D Gaussian Splatting.
Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim process must recover scene geometries and object states, infer physical parameters, and assemble actors, objects, cameras, poses, and trajectories into a runnable physical simulation. Today this process still depends on brittle workflow glue across visual perception tools and simulators: manual tuning of visual foundation models, mesh cleanup, coordinate frame alignments, etc. We introduce \textit{Agentic Real2Sim}, a framework for generalized physical world modeling with vision-language agents that converts a real-world recording of object-robot interaction into a simulatable episodic twin, and connects the resulting twin to downstream policy fine-tuning and evaluation. We evaluate Agentic Real2Sim on rigid-object manipulation, deformable-object interaction, and humanoid motion scenes, spanning domains that are usually handled by separate Real2Sim pipelines. The framework's agentic decisions can be driven by an open-weight VLM backend at a small fraction of the cost of frontier models, while attaining a comparable conversion success rate. The framework further supports custom scene conversion, fine-tuning of a pretrained policy with data generated from converted episodes, and works effectively as a surrogate for real-world policy evaluation. The project site, including code is available at https://agentic-real2sim.github.io.
Guanxiong Chen, Qianjun Xia, Jiawei Peng +24
University of British Columbia · National University of Singapore · Johns Hopkins University +4
Learning physically plausible dynamics from visual observations is essential for interactive world models and embodied agents. However, modeling real-world deformable objects remains challenging because their dynamics often arise from complex, spatially heterogeneous material responses. To address this challenge, we propose PhysReal, a video-driven framework for learning and simulating the underlying physics of real deformable objects. PhysReal integrates a spatially varying hybrid expert-neural constitutive model with a differentiable MPM simulator and 3DGS renderer. Analytical expert models provide interpretable physical priors, while neural constitutive residuals capture material responses beyond predefined formulations. Spatially distributed patches parameterize the constitutive field, enabling a continuous representation of local material variations. To organize the identification of this model from sparse visual observations, we adopt a progressive curriculum that sequentially optimizes global material properties, spatially varying local parameters, and neural constitutive residuals, together with complementary motion and mask supervision. Extensive experiments on diverse deformable-object interactions demonstrate that PhysReal achieves superior performance in dynamic reconstruction and future-state prediction, while showing strong potential for downstream robotic applications.