From Solo to Ensemble: A Hierarchical Framework for Composable Multi-Agent Human-Object Interaction
Organizations: ShanghaiTech University · InstAdapt
Abstract
Physics-based human-object interaction has achieved robust single-agent manipulation skills, yet extending them to multi-agent cooperative tasks remains challenging. Existing approaches typically adapt interaction policies through task-specific fine-tuning, which entangles low-level contact-rich execution with high-level coordination and limits reuse across object geometries, interaction types, and team sizes. We propose a hierarchical framework that converts a single-agent HOI policy into a reusable Object-oriented Motion Skill. Specifically, we reinterpret teacher rollouts as object-oriented action supervision by extracting short-horizon object-proxy motions from executed trajectories, and distill task-specific teachers into a low-level skill operating in an Object-oriented Action Space. For downstream tasks, the distilled skill is frozen as a reusable executor, while a high-level policy coordinates multiple agents by generating region-wise object-oriented actions conditioned on the shared object, task goal, agent states, and local manipulation regions. This formulation shifts multi-agent HOI learning from direct contact-rich full-body control to compact object-level proxy-motion coordination. Experiments on diverse HOI tasks show that the distilled Object-oriented Motion Skill supports robust proxy-motion execution and enables composable policy learning across different interaction types, object geometries, and team sizes.
Figures & tables
| Task | Object | #Agents | CooHOI | Ours | ||
|---|---|---|---|---|---|---|
| Succ. | Precision | Succ. | Precision | |||
| Carrying | Armchair | 1 | 97.26 | 5.05 | 97.81 | 4.99 |
| Table | 1 | 97.07 | 5.23 | 97.19 | 5.11 | |
| Sofa | 2 | 84.17 | 10.12 | 88.21 | 8.70 | |
| H-box | 4 | 80.96 | 17.50 | 85.63 | 14.24 | |
| Cross-box | 4 | 82.38 | 15.13 | 86.14 | 13.81 | |
| Method | Teacher-Rollout Targets | Random-Motion Targets | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| TTR | Succ traj | TTR | Succ traj | |||||||
| InterPhys [ 8 ] | 106.5 | 47.8 | 50.9 | 82.6 | 68.2 | 112.2 | 63.9 | 54.1 | 76.4 | 60.8 |
| SkillMimic [ 63 ] | 91.0 | 40.3 | 44.2 | 90.8 | 81.2 | 99.3 | 54.6 | 49.7 | 84.9 | 68.3 |
| Ours | 79.8 | 38.2 | 31.3 | 100.0 | 98.9 | 95.8 | 45.2 | 39.2 | 96.3 | 91.8 |
| Ours w/o Online Distill. | 117.9 | 52.3 | 50.8 | 86.7 | 84.8 | 148.2 | 63.1 | 55.9 | 58.4 | 23.8 |
| Ours w/o Random-Motion RL | 88.1 | 41.3 | 32.8 | 97.2 | 98.1 | 156.1 | 65.5 | 57.2 | 63.1 | 38.7 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Item | Value |
|---|---|
| Simulator | Isaac Gym |
| Number of environments | |
| Humanoid model | AMP humanoid |
| Control mode | PD target control |
| Control frequency | Hz |
| Physics substeps |