Text-conditioned full-body human-object interaction (HOI) generation requires synthesizing human motion and object trajectories that match the input text while remaining precisely coordinated over time. Most methods represent the human and object as separate trajectories and predict the global human-object couplings. Learning this complex, dynamically changing relationship implicitly, however, often yields object drift, missed contact, and penetration. We introduce PAMI, a Part-Anchored Motion framework for Interaction generation. Inspired by the classic Hough Transform, our key idea is to localize object motion by letting body-part anchors vote for it: we express object motion relative to multiple body-part anchors and use PamiVAE to learn an interaction latent space, decoding frame-wise weights that aggregate these part-specific votes. Building on this representation, PAMI generates interactions in a coarse-to-fine hierarchy. PamiGen first generates a coarse human-object interaction from text in this structured latent space, and PamiRefiner then recursively resolves fine-grained contact geometry using a hybrid surface-sensing representation, combining long-range probes that capture overall body-part influence with short-range sensors that resolve detailed contacts near the object surface. Experiments on InterAct show that PAMI generates more faithful interactions and more accurate human-relative object motion than previous methods, achieving 14.5% higher contact recall than the previous state of the art. Extensive ablations validate the contributions of both the part-anchored voting representation and hybrid surface-sensing refinement.
Figures & tables
Figure 1: Text-conditioned full-body HOI generation. Given a text prompt and a canonical object mesh, PAMI generates coordinated human and object motion across diverse one-handed, two-handed, and full-body interactions. Insets highlight fine-grained contact.
Figure 2: Overview of PAMI. Body-part anchors vote for the object motion, and PamiVAE embeds this part-anchored interaction into a latent space ( Sec. 3.1 ). Generation then proceeds coarse-to-fine: PamiGen samples a coarse interaction from text in the latent space ( Sec. 3.2 ), which PamiRefiner refines in explicit space using hybrid long- and short-range surface-sensing probes to sharpen contacts ( Sec. 3.3 ).
Method
R-Precision (Top 1) ↑
FID ↓
Multimodal Dist. ↓
FSR ↓
Pene ↓
Contact →
Interaction (hands) ↑
Cprec
Crec
CF1
Ground Truth
0.859 ±0.000
0.000 ±0.000
1.475 ±0.003
0.051 ±0.000
0.060 ±0.000
0.219 ±0.000
1.000 ±0.000
1.000 ±0.000
1.000 ±0.000
HOI-Diff
0.703 ±0.015
0.698 ±0.017
2.906 ±0.002
0.042 ±0.000
0.092 ±0.004
0.090 ±0.001
0.754 ±0.009
0.531 ±0.007
0.574 ±0.009
CHOIS
0.717 ±0.005
0.544 ±0.003
2.659 ±0.006
0.072 ±0.001
0.122 ±0.005
0.141 ±0.002
0.823 ±0.003
0.701 ±0.007
0.723 ±0.004
InterDiff
0.779 ±0.000
0.302 ±0.018
2.364 ±0.004
0.069 ±0.001
0.110 ±0.006
0.150 ±0.001
0.834 ±0.008
0.718 ±0.003
0.741 ±0.004
Text2HOI
0.710 ±0.007
0.284 ±0.012
2.551 ±0.011
0.035 ±0.001
0.104 ±0.002
0.108 ±0.000
0.840 ±0.010
0.617 ±0.007
0.670 ±0.006
Table 1: Quantitative evaluation on InterAct ( Xu et al., 2025a ) . R-precision is evaluated with a batch size of 64. → indicates that closer to the ground truth is better. ± indicates a 95% confidence interval. Our method achieves best performance on almost all metrics.
R-Precision (Top 1) ↑
FID ↓
Pene ↓
Contact →
Interaction (full body) ↑
Interaction (hands) ↑
Cprec
Crec
CF1
Cprec
Crec
CF1
Ground Truth
0.859
0.000
0.060
0.219
1.000
1.000
1.000
1.000
1.000
1.000
Our full PAMI model
0.840 ±0.003
0.092 ±0.003
0.062 ±0.001
0.193 ±0.001
0.502 ±0.003
0.477 ±0.003
0.457 ±0.004
0.843 ±0.002
0.836 ±0.006
0.821 ±0.004
(a) w/o part factorization
0.822 ±0.002
0.171 ±0.006
0.086 ±0.001
0.129 ±0.002
0.466 ±0.002
0.294 ±0.005
0.327 ±0.004
0.823 ±0.007
0.657 ±0.006
0.697 ±0.006
(b) w/o leg and root anchors
0.840 ±0.002
0.097 ±0.006
0.087 ±0.002
0.174 ±0.001
0.499 ±0.005
0.425 ±0.004
0.425 ±0.004
0.849 ±0.002
0.806 ±0.004
0.804 ±0.003
(c) w/o anchors
0.834 ±0.003
0.102 ±0.004
0.076 ±0.003
0.116 ±0.001
0.478 ±0.015
0.277 ±0.007
0.313 ±0.009
0.830 ±0.006
0.647 ±0.003
0.686 ±0.001
Table 2: Ablation for part anchored voting representation . Our part factorization, anchor design with learned routing and absolute root representation all contribute to better generation.
R-Precision (Top 1) ↑
FID ↓
Pene ↓
Contact →
Interaction (full body) ↑
Interaction (hands) ↑
Cprec
Crec
CF1
Cprec
Crec
CF1
Ground Truth
0.859
0.000
0.060
0.219
1.000
1.000
1.000
1.000
1.000
1.000
(a) no refinement
0.839 ±0.004
0.103 ±0.004
0.115 ±0.003
0.142 ±0.001
0.500 ±0.006
0.364 ±0.005
0.389 ±0.004
0.849 ±0.001
0.758 ±0.003
0.775 ±0.002
(b) w/o generation stream
0.841 ±0.003
0.101 ±0.005
0.091 ±0.001
0.155 ±0.000
0.496 ±0.003
0.388 ±0.002
0.404 ±0.001
0.847 ±0.002
0.788 ±0.003
0.794 ±0.002
(c) w/o corruption stream
0.841 ±0.003
0.104 ±0.003
0.085 ±0.003
0.166 ±0.001
0.496 ±0.005
0.414 ±0.004
0.419 ±0.002
0.847 ±0.005
0.786 ±0.004
0.793 ±0.004
(d) w/o short-range probes
0.840 ±0.003
0.100 ±0.004
0.083 ±0.002
0.186 ±0.000
0.496 ±0.002
0.462 ±0.001
0.446 ±0.000
0.845 ±0.002
0.828 ±0.003
0.816 ±0.002
Table 3: PamiRefiner ablations . All variants refine the same coarse generations. Mixed data training and hybrid sensing are important to obtain high-quality fine interaction details.
Figure 6
Figure 4: Qualitative comparison. Each row shows four frames from InterAct, LIGHT, and PAMI. InterAct and LIGHT often produce incorrect object orientation or interaction with missing contacts. PAMI better follows the described motion while maintaining coherent contact.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Method
R-Prec. Top 2 ↑
R-Prec. Top 3 ↑
Diversity →
Ground Truth
0.967 ±0.001
0.992 ±0.002
7.764 ±0.020
HOI-Diff
0.881 ±0.012
0.929 ±0.003
7.577 ±0.054
CHOIS
0.880 ±0.004
0.934 ±0.002
7.777 ±0.046
InterDiff
0.930 ±0.002
0.968 ±0.005
7.738 ±0.058
Text2HOI
0.869 ±0.005
0.926 ±0.000
7.730 ±0.009
LIGHT
0.877 ±0.007
0.930 ±0.001
7.737 ±0.062
Appendix
Table 5: Additional metrics on InterAct. Top-2/3 R-precision and Diversity complement Table 1 .
Figure 5: Effect of recursion depth K . From left to right: full-body contact F1, FID, and hand-contact F1. Curves show the mean over three generation draws; bands denote 95% confidence intervals.
Variant
MPJPE (mm) ↓
hand (mm) ↓
object comp. (cm) ↓
object pos. (cm) ↓
object rot. ( ∘ ) ↓
part-factored (ours)
15.2
1.3
2.54
1.66
2.2
+ contact-tuned decoder (final)
16.7
1.3
2.49
1.26
1.6
whole-body, param.-matched
38.8
2.3
2.70
1.38
2.0
wrist-only anchors
16.5
1.3
2.59
1.02
2.2
no anchors
17.7
1.2
–
1.41
1.2
w/o spatiotemporal attention
15.4
1.2
2.94
2.10
2.5
Appendix
Table 6: PamiVAE reconstruction. Validation errors from the final checkpoint of each variant. Body and hand errors are in mm; object translation and rotation errors are in cm and degrees.
Top-1 ↑
Top-2 ↑
Top-3 ↑
FID ↓
MM Dist ↓
Diversity →
FSR ↓
Ground Truth
0.859
0.967
0.992
0.000
1.475
7.764
0.051
Our full PAMI model
0.839 ±0.004
0.957 ±0.002
0.980 ±0.002
0.103 ±0.004
2.198 ±0.015
7.683 ±0.008
0.137 ±0.002
(a) w/o part factorization
0.816 ±0.002
0.946 ±0.001
0.972 ±0.002
0.330 ±0.009
2.375 ±0.004
7.535 ±0.016
0.262 ±0.001
(b) w/o leg and root anchors
0.840 ±0.001
0.960 ±0.003
0.984 ±0.001
0.109 ±0.006
2.215 ±0.014
7.706 ±0.006
0.137 ±0.002
(c) w/o anchors
0.832 ±0.004
0.953 ±0.003
0.977 ±0.001
0.119 ±0.007
2.230 ±0.004
7.723 ±0.013
0.128 ±0.002
(d) w/o learned voting weights
0.830 ±0.003
0.950 ±0.001
0.974 ±0.001
0.116 ±0.006
2.237 ±0.007
7.689 ±0.017
0.135 ±0.001
Appendix
Table 7: Generator ablations before refinement ( K=0 ). These coarse outputs isolate the representation choices from geometric refinement. We report the mean and 95% confidence interval over three generation draws.
Top-2 ↑
Top-3 ↑
MM Dist ↓
Diversity →
FSR ↓
Ground Truth
0.967
0.992
1.475
7.764
0.051
Our full PAMI model
0.958 ±0.002
0.980 ±0.001
2.195 ±0.016
7.693 ±0.007
0.078 ±0.001
(a) w/o part factorization
0.948 ±0.000
0.973 ±0.001
2.308 ±0.003
7.651 ±0.019
0.115 ±0.001
(b) w/o leg and root anchors
0.960 ±0.002
0.985 ±0.001
2.207 ±0.013
7.718 ±0.004
0.112 ±0.001
(c) w/o anchors
0.953 ±0.002
0.976 ±0.001
2.218 ±0.003
7.727 ±0.010
0.092 ±0.001
(d) w/o learned voting weights
0.950 ±0.002
0.974 ±0.001
2.228 ±0.006
7.695 ±0.014
0.103 ±0.001
Appendix
Table 8: Additional metrics for the refined representation ablations in Table 2 ( K=4 ). Mean and 95% confidence interval over three runs.
Top-1 ↑
Top-2 ↑
Top-3 ↑
FID ↓
MM Dist ↓
Diversity →
FSR ↓
Pene ↓
Contact →
Interaction (full body) ↑
Interaction (hands) ↑
Cprec
Crec
CF1
Cprec
Crec
CF1
Ground Truth
0.859
0.967
0.992
0.000
1.475
7.764
0.051
0.060
0.219
1.000
1.000
1.000
1.000
1.000
1.000
LIGHT
0.763 ±0.013
0.911 ±0.009
0.948 ±0.004
0.225 ±0.073
2.643 ±0.013
7.593 ±0.012
0.105 ±0.003
0.134 ±0.005
0.128 ±0.000
0.468 ±0.001
0.308 ±0.004
0.328 ±0.000
0.808 ±0.003
0.686 ±0.013
0.704 ±0.010
LIGHT + PamiRefiner w/o corruption stream
0.762 ±0.009
0.910 ±0.009
0.948 ±0.002
0.241 ±0.068
2.644 ±0.011
7.606 ±0.022
0.061 ±0.001
0.094 ±0.003
0.151 ±0.000
0.459 ±0.001
0.360 ±0.006
0.361 ±0.001
0.795 ±0.002
0.704 ±0.011
0.712 ±0.008
LIGHT + PamiRefiner
0.765 ±0.010
0.912 ±0.010
0.948 ±0.006
0.235 ±0.076
2.642 ±0.009
7.603 ±0.016
0.059 ±0.002
0.069 ±0.003
0.181 ±0.000
0.476 ±0.000
0.429 ±0.001
0.408 ±0.002
0.808 ±0.003
0.765 ±0.009
0.755 ±0.006
InterAct
0.803 ±0.010
0.928 ±0.004
0.962 ±0.005
0.328 ±0.013
2.524 ±0.037
7.827 ±0.047
0.101 ±0.002
0.126 ±0.005
0.122 ±0.001
0.418 ±0.007
0.249 ±0.005
0.278 ±0.001
0.777 ±0.003
0.631 ±0.000
0.658 ±0.002
Appendix
Table 9: Full metrics for Table 4 . Zero-shot refinement of LIGHT and InterAct outputs with K=4 ; mean and 95% confidence interval over two sampling campaigns. Bold marks the best result in each generator block.
Top-2 ↑
Top-3 ↑
MM Dist ↓
Diversity →
FSR ↓
Ground Truth
0.967
0.992
1.475
7.764
0.051
(a) no refinement
0.957 ±0.002
0.980 ±0.002
2.198 ±0.015
7.683 ±0.008
0.137 ±0.002
(b) w/o generation stream
0.958 ±0.002
0.980 ±0.002
2.208 ±0.015
7.696 ±0.009
0.134 ±0.001
(c) w/o corruption stream
0.958 ±0.002
0.980 ±0.003
2.190 ±0.014
7.692 ±0.008
0.085 ±0.000
(d) w/o short-range probes
0.958 ±0.002
0.980 ±0.002
2.198 ±0.014
7.691 ±0.006
0.107 ±0.001
(e) w/o long-range probes
0.958 ±0.003
0.980 ±0.003
2.202 ±0.016
7.683 ±0.007
0.105 ±0.002
Appendix
Table 10: Additional metrics for the PamiRefiner ablations in Table 3 ( K=4 ). Mean and 95% confidence interval over three runs.
Text driven hand object interaction (HOI) generation is gaining attention for immersive applications and robotics, yet producing physically plausible interactions remains challenging. Even when individual motions appear natural, small contact errors can cause conspicuous artifacts such as floating and interpenetration. Prior methods mitigate these issues using explicit contact cues or implicit grasp priors, but typically rely on multi stage pipelines and fail to model temporally evolving contact. We present JointHOI, a single stage diffusion framework that jointly generates 3D hand object motion and dynamic, distance based contact maps from text. By treating contact as an auxiliary inner modality, joint generation enables the model to learn contact motion coupling during training. At inference, contact guided sampling enforces consistency between generated contact maps and motion implied geometry, improving temporal stability and reducing penetration and floating. Experiments on GRAB and ARCTIC demonstrate consistent improvements in text adherence and physical plausibility over prior methods.
Mingyeong Song, Jungbin Cho, Jisoo Kim +5
Ewha Womans University, Seoul, Korea · Yonsei University, Seoul, Korea · Carnegie Mellon University, Pittsburgh PA, USA +1
Generating realistic 3D Human-Object Interactions (HOI) is a fundamental task for applications ranging from embodied AI to virtual content creation, which requires harmonizing high-level semantic intent with strict low-level physical constraints. Existing methods excel at semantic alignment, however, they struggle to maintain precise object contact. We reveal a key finding termed \textit{Geometric Forgetting}: as diffusion model depth increases, semantic feature tend to overshadow object geometry feature, causing the model to lose its perception to object geometry. To address this, we propose MaMi-HOI, a hierarchical framework reconciling \textbf{Ma}cro-level kinematic fluidity with \textbf{Mi}cro-level spatial precision. First, to counteract geometric forgetting, we introduce the Geometry-Aware Proximity Adapter (GAPA), which explicitly re-injects dense object details to perform residual snapping corrections for precise contact. Nevertheless, such aggressive local enforcement can disrupt global dynamics, leading to robotic stiffness. In response, we introduce the Kinematic Harmony Adapter (KHA), which proactively aligns whole-body posture with spatial objectives, ensuring the skeleton actively accommodates constraints without compromising naturalness. Extensive experiments validate that MaMi-HOI simultaneously achieves natural motion and precise contact. Crucially, it extends generation capabilities to long-term tasks with complex trajectories, effectively bridging the gap between global navigation and high-fidelity manipulation in 3D scenes. Code is available at https://github.com/DON738110198/MaMi-HOI.git
Hao Wang, Shiqi Wang, Qi Liu
School of Future Technology, South China University of Technology, Guangzhou, Guangdong, China.
Daily activities require humans to coordinate whole-body motion with the motion of surrounding objects. Despite recent progress in human-object interaction (HOI) generation, most existing methods assume interactions with a single rigid object and do not extend well to scenarios involving a variable number of objects or articulated objects with diverse joint mechanisms. We propose surface keypoint trajectories as an object motion representation: for each rigid component, whether a standalone object or one part of an articulated assembly, we track a small set of non-collinear surface points over time. This representation handles multi-object coordination and diverse articulation mechanisms directly from point dynamics without requiring explicit joint-type specification. To model when and where each body region contacts each object, we introduce a spatio-temporal contact distance field that extends distance-based contact modeling to whole-body, multi-object, and articulated settings. We factorize HOI generation into three stages: generating object motions from text or waypoints, predicting the contact distance field, and synthesizing whole-body motion with contact-guided optimization. Experiments on ParaHome, HIMO, ARCTIC, and OMOMO demonstrate better or comparable performance to existing methods across single-object, multi-object, and articulated interaction settings.
Xiaogang Peng, Zeyu Han, Zichong Meng +4
Northeastern University, USA · Xi’an Jiaotong University, China · Amazon, USA