RiCo: Neural Simulation of Rigid-Body Interactions via Local Contact Reasoning
Authors: Ruixiang Ouyang, Guanren Qiao, Fansen Meng, Yueci Deng, Ruixing Jin, Kui Jia, Guiliang Liu
Organizations: The Chinese University of Hong Kong, Shenzhen · DexForce Co., Ltd. · South China University of Technology · Shenzhen Loop Area Institute
Accurate simulation of rigid-body interactions is essential for predictive physical world models. Despite recent progress in modeling object dynamics, capturing how local contacts between surfaces shape object motion remains challenging. While end-to-end world models predict interactions across entire scenes or objects, in practice, rigid-body contact is inherently local, and only nearby surfaces can directly exchange contact forces. Motivated by this observation, we introduce Rigid-body Contact Reasoning (RiCo), which represents interactions between objects through sparse neighborhoods of contact surface points. RiCo combines each point's state with the relative geometry, motion, and physical properties of nearby surfaces, then reasons across the object's points to determine how these local contacts jointly affect its motion. By confining cross-object reasoning to nearby surfaces while propagating contact information within each rigid body, RiCo retains fine-grained interaction details without the cost of modeling every pair of scene points. Such properties enable RiCo a higher accuracy and contact fidelity. Experiments on MOVi-benchmark demonstrate that RiCo reduces 100-frame position and orientation errors by 31-35% and approximately 38%, respectively, compared with baselines. Moreover, RiCo achieves high contact fidelity, with ground-truth-relative penetration-time and mean-depth differences of 11.0% and 2.22 mm, respectively. RiCo further generalizes zero-shot from small-scale training scenarios to scenes containing 270 objects. Our real-world multi-ball collision experiments further provide preliminary evidence of sim-to-real transfer.
Figures & tables
Figure 1: Comparing rigid body contact modeling.
Figure 2: Overview of RiCo. From two consecutive scene states, RiCo constructs sparse local contact neighborhoods, fuses them with point-wise object states, and performs shared intra-object reasoning with a point transformer. The resulting contact-conditioned features are decoded at a small set of anchors to recover a rigid transformation for each object, and the predicted state is recursively used for autoregressive rollout.
Model
MOVi-A
MOVi-B
MOVi-Sphere
50
75
100
50
75
100
50
75
100
FIGNet reimpl
0.132/7.10
0.285/14.62
0.492/23.30
0.141/7.39
0.300/15.16
0.516/24.96
N/A
N/A
N/A
HCMT reimpl
0.239/5.70
0.538/11.82
0.951/18.40
0.237/4.72
0.527/9.80
0.932/17.43
0.243/4.19
0.541/8.21
0.956/13.81
VPD reimpl
0.235/5.10
0.489/11.66
0.827/20.37
0.275/4.65
0.581/9.70
0.987/16.99
0.244/4.47
0.510/9.50
0.855/17.65
HopNet
0.054/5.64
0.115/11.84
0.196/18.83
0.047 /4.91
0.101/10.35
0.176/17.91
0.034 /4.05
0.073/8.21
0.124/13.68
RigidFormer
0.049 / 5.06
0.103 / 10.90
0.177 / 18.32
0.050/ 3.97
0.095 / 8.51
0.161 / 15.33
0.026 / 3.00
0.057 / 6.48
0.099 / 11.19
Table 1: Performance on MOVi-A, MOVi-B, and MOVi-Sphere. Each cell reports position RMSE (m) / orientation RMSE ( ∘ ) at prediction horizons of 50, 75, and 100 frames.
Train
Test
HopNet
RigidFormer
RiCo
75
100
75
100
75
100
MOVi-S
MOVi-A
0.112/11.29
0.197/ 17.74
0.107 / 11.27
0.183 /17.92
0.110 / 10.15
0.190 / 16.79
MOVi-B
0.106/9.75
0.188/17.13
0.096 / 9.42
0.161 / 16.81
0.103 / 9.63
0.170 / 16.52
MOVi-A
MOVi-S
0.100/9.16
0.172/15.17
0.087 / 7.94
0.153 / 14.27
0.055 / 5.63
0.094 / 9.83
MOVi-B
0.117 /10.77
0.202 /18.66
0.123/ 9.29
0.208/ 17.08
0.095 / 10.56
0.164 / 17.93
MOVi-B
MOVi-S
0.095 / 9.09
0.160 / 15.04
0.123/10.52
0.207/17.88
0.049 / 5.31
0.085 / 9.29
Table 2: Generalization performance across MOVi-Sphere (S), MOVi-A (A), and MOVi-B (B). Each reports position RMSE (m) / orientation RMSE ( ∘ ) at prediction horizons of 75 and 100 frames.
Figure 3: Ablation of local contact reasoning on MOVi-B. (a) Ground-truth-relative penetration statistics. (b) Position and orientation RMSE at the 100-step prediction horizon. For each method, the left bar shows the time ratio (left y-axis), and the right bar shows penetration depth (right y-axis).
Figure 4: Resolution and efficiency analysis on MOVi-B. (a) Ground-truth-relative penetration statistics across different point-cloud resolutions. (b) Position and orientation prediction errors. For these curves, solid lines denote time ratios, measured on the left y-axis. Dashed lines denote penetration depths, measured on the right y-axis.
Position / Orientation RMSE
Runtime
Scene
120
240
360
FPS
Spheres
0.029/1.73
0.243/43.07
0.713/89.90
5.02
Spots
0.019/1.21
0.211/22.90
0.785/86.74
6.50
Knots
0.017/0.81
0.139/13.43
0.843/78.00
6.54
Table 3: Large-scale prediction.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Descriptor
Channels (in implementation order)
Dim.
Source-point descriptor st,ni
xt,ni−cti
3
Δxt,ni
3
νt,ni
3
(mi)−1,μi,ei
3
xt,ni−x0,ni
3
Appendix
Table 4: Input feature specification. All geometric positions and finite-difference displacements are expressed in the same scene coordinate system. Here y is an external candidate point, next its unit normal, and Δy its finite-difference displacement.
Component
Configuration
Maximum object slots / points per object
10 / 1024
Anchors / contact candidates per point
8 / 4
Source feature MLP
15→128→384 , GELU
Contact feature MLP
14→128→384 , Softplus
Contact pooling
Masked distance softmax
Contact temperature (initial/minimum)
0.005 / 10−4
Appendix
Table 5: Default local-contact model architecture. The two Transformer stacks operate within each object; the indicated cross-attention is from anchor queries to the object’s own point tokens.
Setting
Value
Optimizer
AdamW
Global batch size
128
Per-device micro-batch size
8
Context / target frames
2 / 1
Precision
bf16
Initial / final learning rate
10−4 / 10−5
Appendix
Table 6: Reference offline optimization settings for the local-contact model.
Figure 5: Runtime analysis.
Figure 6: Qualitative results of three large-scale scenarios.
Figure 7: Temporal error accumulation in large-scale scenes. Position and orientation errors throughout autoregressive rollouts of three 270-object scenarios. The dashed line indicates an approximate transition point after which object collisions become more frequent.
Figure 8: Illustrating 10 pairs qualitative comparison of real-world observations (upper, labeled by Real Obs.) and autoregressive predictions by RiCo (lower, labeled by RiCo Pred.) in three- and four-ball collision scenarios. RiCo’s predictions are initialized by the balls’ observed positions and estimated velocities in the real-world experiments, then replayed in the simulated environment. Under this setting, smaller distances between corresponding predicted and observed positions indicate greater prediction accuracy.
NLPR, MAIS, Institute of Automation, Chinese Academy of Sciences · School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences +1