cs.CVSep 26, 2026

Harnessing Coupled Stream Completion For Human-Object Interaction Modeling

Authors: Dawei Guan, Di Yang, Jiangtao Wang

Organizations: School of Artificial Intelligence and Data Science, University of Science and Technology of China

Abstract

Text-conditioned human-object interaction (HOI) generation requires body motion, object trajectories & rotations, and hand articulation to remain coordinated. These components differ in scale and dynamics, but must agree on contact, relative pose, and timing. A shared representation may limit the distinct structure of each stream, while independent generation prevents each stream from responding to changes in the others. Latent supervision alone also does not directly constrain contact after decoding. We propose TRACE, a continuous latent framework that keeps stream states separate and couples their updates. TRACE encodes body, object, and hand motion into separate latents and predicts each stream velocity from the complete current interaction state. Geometric losses on decoded motion further constrain contact and object-relative motion over time. The same model supports completion of any single absent stream from the other two. Frozen flow features also serve as input to a language model for HOI understanding. Experiments on InterAct, OMOMO, and BEHAVE show that joint completion training improves generation and that frozen flow features improve understanding over raw-motion encoding. On InterAct, TRACE achieves the highest contact precision, recall, and F1 among the compared methods.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Uni-HOI:A Unified framework for Learning the Joint distribution of Text and Human-Object Interaction

    Date pendingMengfei Zhang, Jinlu Zhang, Zhigang TuArticulated ObjectsHuman Motion Prediction

  2. JointHOI: Jointly Generating Contact Maps Enhances Hand Object Interaction Generation

    Jul 2, 2026Mingyeong Song, Jungbin Cho, Jisoo Kim +5Hand-Object Interaction DetectionHuman Motion Generation

  3. PAMI: Part Anchored Motion for Text to Human-Object Interaction Generation

    Sep 29, 2026Chuqiao Li, Xianghui Xie, Yong Cao +2Interaction DataText Analysis and Detection