Organizations: NLPR, MAIS, Institute of Automation, Chinese Academy of Sciences · School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Independent Researcher
Deformable-object manipulation is essential for robotic tasks such as folding laundry and handling food, where robots must control shape changes as well as object motion. Predictive soft-body simulation supports these tasks by anticipating deformation under external interactions. However, spatial neighborhoods can misrepresent deformation dependencies, introducing local errors that accumulate over successive predictions. Models fitted to individual scenes must also accommodate changes in object geometry and manipulation conditions. In this work, we propose AIM, an Adaptive Interaction Modeling framework that treats real-to-sim soft-body simulation as a local-global interaction modeling problem. AIM uses motion history and geometry to adapt particle relations over current spatial neighbors and retained connections, while geometry-conditioned global communication coordinates object-wide responses. A unified kinematic control-point interface represents different manipulation configurations, and multi-step autoregressive supervision trains the model on its own predicted trajectories. Experiments on PhysTwin and PGND demonstrate improved motion accuracy and visual fidelity, with a 20.0% reduction in future-prediction tracking error relative to PhysTwin and a 22.8% reduction in mean long-horizon particle error across six object categories relative to PGND. The framework further supports transfer across actions, object instances, and scenes, including zero-shot transfer from robot interactions to human manipulation without target-domain dynamics fitting.
Figures & tables
Figure 1: Challenges in real-to-sim deformable dynamics.
Figure 2: Overview of AIM . Adaptive particle relations and geometry-conditioned global communication jointly predict object dynamics under kinematic controls.
Reconstruction & Resimulation
Future Prediction
Method
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
PhysTwin
0.005
0.010
84.4
28.636
0.949
0.030
0.012
0.025
68.2
25.321
0.942
0.056
SoMA
0.018
0.058
67.0
25.007
0.917
0.100
0.041
0.148
39.5
21.776
0.865
0.136
PGND
0.009
0.013
82.1
28.228
0.949
0.033
0.024
0.044
60.3
24.426
0.938
0.069
AIM
0.006
0.007
87.0
29.473
0.952
0.025
0.014
0.020
69.8
25.916
0.943
0.055
Table 1: Quantitative comparison on reconstruction & resimulation and future prediction on the PhysTwin benchmark. Best results are shown in bold , and second-best results are underlined .
Figure 3: Qualitative Results on Reconstruction & Resimulation and Future Prediction.
Figure 4: Qualitative Results on Held-out Interaction Prediction.
Figure 5: Long-Horizon Rollout Performance on Held-out Interaction Prediction.
Cross-Action Transfer
Cross-Object Transfer
Method
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
PhysTwin
0.035
0.053
53.1
21.679
0.900
0.110
–
–
–
–
–
–
PGND
0.226
0.274
58.4
23.763
0.902
0.089
0.026
0.038
72.2
23.023
0.934
0.086
AIM
0.020
0.045
70.6
24.430
0.904
0.073
0.018
0.023
79.8
24.506
0.942
0.073
Table 2: Quantitative comparison on cross-action transfer and cross-object transfer on the PhysTwin benchmark. Best results are shown in bold . “–” indicates that transferring PhysTwin’s object-specific spring-mass model requires additional adaptation beyond our evaluation protocol.
Held-Out Scene Generalization
Held-Out Action Generalization: Lift → Fold
Method
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
PGND
0.023
–
84.3
24.118
0.923
0.090
0.032
–
69.1
22.037
0.900
0.109
AIM
0.013
–
87.1
24.869
0.929
0.074
0.019
–
71.3
22.982
0.896
0.108
Table 3: Generalization to held-out cloth scenes and unseen action types . Best results are shown in bold . “–” indicates unavailable ground-truth 3D tracking annotations.
Method
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
PGND
0.019
0.033
60.28
24.835
0.942
0.070
AIM
0.016
0.023
65.21
25.438
0.941
0.069
Table 4: Zero-shot transfer from PGND to PhysTwin. Best results are shown in bold .
Figure 6: Ablation study on double_stretch_zebra
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Additional qualitative comparisons on PhysTwin for reconstruction and resimulation (left) and future prediction (right). Red boxes highlight local deformation differences.
Object
Method
MDE ↓
CD ↓
EMD ↓
IoU ↑
Boundary F ↑
PSNR ↑
SSIM ↑
LPIPS ↓
PGND
45.80
43.07
22.05
65.40
44.22
23.86
0.916
0.108
Cloth
AIM
44.15
38.67
18.45
66.13
43.95
23.92
0.919
0.106
PGND
47.52
47.47
26.35
25.43
46.68
23.34
0.975
0.042
Rope
AIM
22.49
20.63
10.87
32.24
57.50
23.39
0.977
0.034
PGND
45.07
34.36
18.03
58.90
47.20
23.40
0.932
0.069
Plush
AIM
38.71
29.89
14.70
58.62
50.30
23.44
0.935
0.067
Appendix
Table 5: Quantitative comparison with PGND on unseen interaction prediction across six deformable object categories. Best results are shown in bold .
Figure 8: Qualitative Results on Cross-Action Transfer on the Same Object.
Method
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Cloth 3: Double lift → Single clift
PhysTwin
0.010
0.029
88.6
24.310
0.877
0.126
PGND
0.011
0.023
94.9
27.932
0.882
0.100
AIM
0.011
0.022
95.3
28.543
0.886
0.047
Sloth: Double lift → Double stretch
PhysTwin
0.051
0.063
35.4
19.662
0.899
0.132
Appendix
Table 6: Per-scene results on cross-action transfer. Each model is transferred from the source interaction to the target interaction without target-specific training. Best results within each transfer pair are shown in bold .
Figure 9: Qualitative Results on Cross-Object Transfer under Matched Action Types.
Method
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Single lift: Cloth 1 → Cloth 4
PGND
0.040
0.053
62.4
20.332
0.922
0.107
AIM
0.023
0.029
76.9
23.066
0.939
0.081
Double lift: Cloth 3 → Cloth 1
PGND
0.012
0.021
81.9
25.715
0.945
0.065
AIM
0.012
0.016
82.6
25.945
0.944
0.064
Appendix
Table 7: Per-scene results on cross-object transfer under matched action types. Best displayed results within each transfer pair are shown in bold .
Method
CD ↓
Track ↓
IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Rope: robot → human manipulation
PGND
0.014
0.029
39.70
26.132
0.976
0.035
AIM
0.013
0.017
54.00
27.738
0.977
0.030
Cloth: robot → human manipulation
PGND
0.021
0.032
76.60
25.130
0.926
0.088
AIM
0.017
0.022
78.50
25.561
0.927
0.085
Appendix
Table 8: Per-scene results on zero-shot transfer from PGND to PhysTwin. Best results within each scene are shown in bold .