RiCo: Neural Simulation of Rigid-Body Interactions via Local Contact Reasoning
Authors: Ruixiang Ouyang, Guanren Qiao, Fansen Meng, Yueci Deng, Ruixing Jin, Kui Jia, Guiliang Liu
Organizations: The Chinese University of Hong Kong, Shenzhen · DexForce Co., Ltd. · South China University of Technology · Shenzhen Loop Area Institute
Accurate simulation of rigid-body interactions is essential for predictive physical world models. Despite recent progress in modeling object dynamics, capturing how local contacts between surfaces shape object motion remains challenging. While end-to-end world models predict interactions across entire scenes or objects, in practice, rigid-body contact is inherently local, and only nearby surfaces can directly exchange contact forces. Motivated by this observation, we introduce Rigid-body Contact Reasoning (RiCo), which represents interactions between objects through sparse neighborhoods of contact surface points. RiCo combines each point's state with the relative geometry, motion, and physical properties of nearby surfaces, then reasons across the object's points to determine how these local contacts jointly affect its motion. By confining cross-object reasoning to nearby surfaces while propagating contact information within each rigid body, RiCo retains fine-grained interaction details without the cost of modeling every pair of scene points. Such properties enable RiCo a higher accuracy and contact fidelity. Experiments on MOVi-benchmark demonstrate that RiCo reduces 100-frame position and orientation errors by 31-35% and approximately 38%, respectively, compared with baselines. Moreover, RiCo achieves high contact fidelity, with ground-truth-relative penetration-time and mean-depth differences of 11.0% and 2.22 mm, respectively. RiCo further generalizes zero-shot from small-scale training scenarios to scenes containing 270 objects. Our real-world multi-ball collision experiments further provide preliminary evidence of sim-to-real transfer.
Figures & tables
Figure 1: Comparing rigid body contact modeling.
Figure 2: Overview of RiCo. From two consecutive scene states, RiCo constructs sparse local contact neighborhoods, fuses them with point-wise object states, and performs shared intra-object reasoning with a point transformer. The resulting contact-conditioned features are decoded at a small set of anchors to recover a rigid transformation for each object, and the predicted state is recursively used for autoregressive rollout.
Model
MOVi-A
MOVi-B
MOVi-Sphere
50
75
100
50
75
100
50
75
100
FIGNet reimpl
0.132/7.10
0.285/14.62
0.492/23.30
0.141/7.39
0.300/15.16
0.516/24.96
N/A
N/A
N/A
HCMT reimpl
0.239/5.70
0.538/11.82
0.951/18.40
0.237/4.72
0.527/9.80
0.932/17.43
0.243/4.19
0.541/8.21
0.956/13.81
VPD reimpl
0.235/5.10
0.489/11.66
0.827/20.37
0.275/4.65
0.581/9.70
0.987/16.99
0.244/4.47
0.510/9.50
0.855/17.65
HopNet
0.054/5.64
0.115/11.84
0.196/18.83
0.047 /4.91
0.101/10.35
0.176/17.91
0.034 /4.05
0.073/8.21
0.124/13.68
RigidFormer
0.049 / 5.06
0.103 / 10.90
0.177 / 18.32
0.050/ 3.97
0.095 / 8.51
0.161 / 15.33
0.026 / 3.00
0.057 / 6.48
0.099 / 11.19
Table 1: Performance on MOVi-A, MOVi-B, and MOVi-Sphere. Each cell reports position RMSE (m) / orientation RMSE ( ∘ ) at prediction horizons of 50, 75, and 100 frames.
Train
Test
HopNet
RigidFormer
RiCo
75
100
75
100
75
100
MOVi-S
MOVi-A
0.112/11.29
0.197/ 17.74
0.107 / 11.27
0.183 /17.92
0.110 / 10.15
0.190 / 16.79
MOVi-B
0.106/9.75
0.188/17.13
0.096 / 9.42
0.161 / 16.81
0.103 / 9.63
0.170 / 16.52
MOVi-A
MOVi-S
0.100/9.16
0.172/15.17
0.087 / 7.94
0.153 / 14.27
0.055 / 5.63
0.094 / 9.83
MOVi-B
0.117 /10.77
0.202 /18.66
0.123/ 9.29
0.208/ 17.08
0.095 / 10.56
0.164 / 17.93
MOVi-B
MOVi-S
0.095 / 9.09
0.160 / 15.04
0.123/10.52
0.207/17.88
0.049 / 5.31
0.085 / 9.29
Table 2: Generalization performance across MOVi-Sphere (S), MOVi-A (A), and MOVi-B (B). Each reports position RMSE (m) / orientation RMSE ( ∘ ) at prediction horizons of 75 and 100 frames.
Figure 3: Ablation of local contact reasoning on MOVi-B. (a) Ground-truth-relative penetration statistics. (b) Position and orientation RMSE at the 100-step prediction horizon. For each method, the left bar shows the time ratio (left y-axis), and the right bar shows penetration depth (right y-axis).
Figure 4: Resolution and efficiency analysis on MOVi-B. (a) Ground-truth-relative penetration statistics across different point-cloud resolutions. (b) Position and orientation prediction errors. For these curves, solid lines denote time ratios, measured on the left y-axis. Dashed lines denote penetration depths, measured on the right y-axis.
Position / Orientation RMSE
Runtime
Scene
120
240
360
FPS
Spheres
0.029/1.73
0.243/43.07
0.713/89.90
5.02
Spots
0.019/1.21
0.211/22.90
0.785/86.74
6.50
Knots
0.017/0.81
0.139/13.43
0.843/78.00
6.54
Table 3: Large-scale prediction.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Descriptor
Channels (in implementation order)
Dim.
Source-point descriptor st,ni
xt,ni−cti
3
Δxt,ni
3
νt,ni
3
(mi)−1,μi,ei
3
xt,ni−x0,ni
3
Appendix
Table 4: Input feature specification. All geometric positions and finite-difference displacements are expressed in the same scene coordinate system. Here y is an external candidate point, next its unit normal, and Δy its finite-difference displacement.
Component
Configuration
Maximum object slots / points per object
10 / 1024
Anchors / contact candidates per point
8 / 4
Source feature MLP
15→128→384 , GELU
Contact feature MLP
14→128→384 , Softplus
Contact pooling
Masked distance softmax
Contact temperature (initial/minimum)
0.005 / 10−4
Appendix
Table 5: Default local-contact model architecture. The two Transformer stacks operate within each object; the indicated cross-attention is from anchor queries to the object’s own point tokens.
Setting
Value
Optimizer
AdamW
Global batch size
128
Per-device micro-batch size
8
Context / target frames
2 / 1
Precision
bf16
Initial / final learning rate
10−4 / 10−5
Appendix
Table 6: Reference offline optimization settings for the local-contact model.
Figure 5: Runtime analysis.
Figure 6: Qualitative results of three large-scale scenarios.
Figure 7: Temporal error accumulation in large-scale scenes. Position and orientation errors throughout autoregressive rollouts of three 270-object scenarios. The dashed line indicates an approximate transition point after which object collisions become more frequent.
Figure 8: Illustrating 10 pairs qualitative comparison of real-world observations (upper, labeled by Real Obs.) and autoregressive predictions by RiCo (lower, labeled by RiCo Pred.) in three- and four-ball collision scenarios. RiCo’s predictions are initialized by the balls’ observed positions and estimated velocities in the real-world experiments, then replayed in the simulated environment. Under this setting, smaller distances between corresponding predicted and observed positions indicate greater prediction accuracy.
Learning-based simulation of multi-object rigid-body dynamics remains difficult because contact is discontinuous and errors compound over long horizons. Most existing methods remain tied to mesh connectivity and vertex-level message passing, which limits their applicability to mesh-free inputs such as point clouds and leads to high computational cost. Efficiently modeling high-fidelity rigid-body dynamics from mesh-free representations, therefore, remains challenging. We introduce RigidFormer, an object-centric Transformer-based model that learns mesh-free rigid-body dynamics with controllable integration step sizes. RigidFormer reasons at the object level and advances each object through compact anchors; Anchor-Vertex Pooling enriches these anchors with local vertex features, retaining contact-relevant geometry without dense vertex-level interaction. We propose Anchor-based RoPE to inject anchor geometry into attention while respecting the unordered nature of objects and anchors: object-token processing is permutation-equivariant, and the mean-pooled anchor descriptor is invariant to anchor reindexing while preserving shape extent. RigidFormer further enforces rigidity by projecting updates onto the rigid-body manifold using differentiable Kabsch alignment. On standard benchmarks, RigidFormer outperforms or matches mesh-based baselines using point inputs, runs faster, generalizes to unseen point resolutions and across datasets, and scales to 200+ objects; we also show a preliminary extension to command-conditioned articulated bodies by treating body parts as interacting object-level components.
Adapting robots to new objects and tasks requires interaction experience that can be costly to obtain. We present WorldContact, a contact-centric world model for deformable-object manipulation, constructed from a limited set of high-quality trajectories to generate additional training data efficiently. It predicts object dynamics using larger time steps than the source numerical simulator, which requires small integration steps to resolve rapid motion and prevent interpenetration. We evaluate WorldContact across 16 shopping-bag manipulation tasks. State-rollout measurements on a single H100 GPU show a 10× speedup over the source simulator, excluding rendering and disk I/O. We use the generated data to fine-tune an existing vision-language-action policy and deploy it directly on a real robot. In bag lifting, the same policy achieves 65% single-attempt success when fine-tuned on source simulation data alone, compared with 95% when fine-tuned on the dataset expanded with WorldContact. These results support efficient data generation with WorldContact for robot policy adaptation.
Caoliwen Wang, Mengdi Wang, Heng Zhang +7
University of British Columbia · Westlake University · Style3D Research
Deformable-object manipulation is essential for robotic tasks such as folding laundry and handling food, where robots must control shape changes as well as object motion. Predictive soft-body simulation supports these tasks by anticipating deformation under external interactions. However, spatial neighborhoods can misrepresent deformation dependencies, introducing local errors that accumulate over successive predictions. Models fitted to individual scenes must also accommodate changes in object geometry and manipulation conditions. In this work, we propose AIM, an Adaptive Interaction Modeling framework that treats real-to-sim soft-body simulation as a local-global interaction modeling problem. AIM uses motion history and geometry to adapt particle relations over current spatial neighbors and retained connections, while geometry-conditioned global communication coordinates object-wide responses. A unified kinematic control-point interface represents different manipulation configurations, and multi-step autoregressive supervision trains the model on its own predicted trajectories. Experiments on PhysTwin and PGND demonstrate improved motion accuracy and visual fidelity, with a 20.0% reduction in future-prediction tracking error relative to PhysTwin and a 22.8% reduction in mean long-horizon particle error across six object categories relative to PGND. The framework further supports transfer across actions, object instances, and scenes, including zero-shot transfer from robot interactions to human manipulation without target-domain dynamics fitting.
Tiancheng Yang, Dingshuo Chen, Tianle Chen +2
NLPR, MAIS, Institute of Automation, Chinese Academy of Sciences · School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences +1