Recent methods have made promising progress in generating interactions between two humanoids, largely relying on physics-based tracking policies to convert digital reference motions into executable trajectories. However, limited tracking capabilities restrict the range of reference motions that can be successfully executed, reducing data utilization. Moreover, even successful tracking does not guarantee physically plausible responses or faithful realization of the intended interactions. In this paper, we introduce DIGHT, a co-adaptive framework that couples a Digital human Interaction Generator with a Humanoid Tracking policy. Our DIGHT first executes multiple text-conditioned interaction candidates in simulation using a fixed tracker. It then constructs physics-grounded preferences from the resulting rollouts, covering both general executability and interaction fidelity. Rather than collapsing these signals into a single scalar reward for candidate ranking, we align the pretrained generator using physics-decoupled diffusion direct preference optimization (DPO), preserving criterion-specific supervision without differentiating through the simulator. To improve executability, preference pairs are derived from tracking error, friction, and floating. Additionally, to improve interaction fidelity, we propose to incorporate force feedback from simulator as a measure of contact fidelity and construct preferences over contact occurrence, location, duration, and force magnitude. The aligned generator then supplies reference motions for fine-tuning the tracker, improving compatibility between generation and physical execution. Extensive experiments demonstrate that our approach not only improves the physical plausibility of generated motions but also enables more reliable and faithful humanoid interactions in simulation.
Figures & tables
Figure 1: Comparison between existing methods and DIGHT. (a) Direct tracking may miss intended contacts or induce excessive forces despite successful pose tracking. (b) InterAgent produces physically implausible contact forces and unstable interactions due to the gap between the training and inference. (c) In contrast, DIGHT uses physical feedback from generated rollouts to co-adapt the interaction generator and tracking policy, preserving intended contacts while producing physically stable interactions.
Figure 2: Overview of DIGHT. DIGHT consists of an interaction generator and a tracking policy, which are co-adapted in two stages. First, multiple interaction candidates produced by the generator are executed by the frozen tracker, whose rollouts are used to construct physics-grounded preferences for tracking, frictional dissipation, floating, contact occurrence, and contact quality. These preferences align the generator through physics-decoupled diffusion DPO. Second, motions from the aligned generator are used to adapt the tracking policy, improving their compatibility.
Figure 3: Qualitative comparison. We compare our DIGHT with InterGen under various text prompts. Our method generates more natural and physically plausible humanoid interactions, with more realistic and coherent physical contacts, while better aligning with the given text.
Figure 4: Qualitative results of long-horizon humanoid interaction generation
Figure 5: Qualitative results of multi-humanoid interaction generation.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Iter.
Succ. (%)
Fall (%)
MBD (cm)
1
92.19
5.77
7.86
2
92.88
5.31
7.32
3
93.05
4.96
7.41
4
92.51
5.48
7.68
Appendix
Table 7: Effect of iterative optimization.
Figure 6: Additional qualitative comparisons between our method and the baseline.
Figure 7: Additional qualitative comparisons between our method and the baseline.
Figure 8: Additional qualitative comparisons between our method and the baseline.
We address the task of generating physically accurate and visually faithful 4D Human-Object Interaction (HOI). Given a static 3D human and target object represented as 3D Gaussian Splats (3DGS), our goal is to synthesize dynamic scenes where the human actively engages with the object through actions, such as punching or kicking, in accordance with a given input text. To this end, we introduce PhyGenHOI, a novel framework that couples generative human motion with an explicit physical object simulation. We model the human as a semantic agent driven by a Motion Diffusion Model (MDM) and the object as a physical agent simulated via the Material Point Method (MPM), utilizing 3D Gaussians as a unified, differentiable representation. We supervise their interaction through three coupled mechanisms: (1) A Windowed Attraction Loss that temporally synchronizes generative motion to intercept the object; (2) A Contact-Driven Re-simulation step that triggers physically consistent momentum transfer upon impact; and (3) A Masked Video-SDS objective that injects video-based priors to enhance contact fidelity. Experiments show PhyGenHOI generates physically consistent 4D HOI across diverse actions, humans, and objects, outperforming baselines. Project page and videos: https://omerbenishu.github.io/PhyGenHOI/
Despite substantial progress in text-driven 3D human motion synthesis, generating realistic multi-person interaction sequences remains challenging. Notably, body inter-penetration is a pervasive issue from both data acquisition to the generated results, which significantly undermines the realism and usability. Previous generative models either ignored this issue or introduced computationally expensive mesh-level loss functions to alleviate inter-body collisions. In this paper, we propose a general-purpose and computationally efficient optimization strategy named PhysiGen to explicitly integrate collision-aware physical constraints for human-human interaction generation. Specifically, we simplify the high-resolution human body mesh into geometric primitives to greatly reduce the cost of inter-person collision detection. Moreover, we identify the collision regions as the guidance of the optimization directions. PhysiGen is plug-and-play and can be readily integrated into existing human interaction generation models. Extensive cross-dataset and cross-model experiments show that our method can effectively reduce interpenetration and significantly improve visual coherence and physical plausibility compared to the state-of-the-art methods.
Nan Lei, Yuan-Ming Li, Ling-An Zeng +5
Sun Yat-sen University · Shanghai Jiao Tong University · Guilin University of Electronic Technology +1
This paper tackles the problem of physics-aware human motion synthesis in a dynamic scene. Unlike existing works which mainly tend to generate physically unrealistic motions due to limited contact modeling, typically restricted to hands, in this paper, we introduce a physics-aware human motion generation framework that explicitly models the full spectrum of human-related forces, including human-object, human-scene, and internal body dynamics.~Our method imposes soft physical constraints to maintain force and torque balance, ensuring physically grounded motion synthesis. We further propose a novel continuous distance-based force model that generalizes contact modeling to arbitrary surfaces, capturing interactions not only with static environments but also with dynamic, moving objects. Extensive experiments show that our approach significantly improves physical plausibility and generalizes well to complex scenes, setting a new benchmark for physically consistent human motion generation.
Chaoyue Xing, Wei Mao, Miaomiao Liu
Australian National University, Canberra, Australia