cs.CVOct 8, 2026

MAMHOI: Factorizing Scene-Aware Human-Object Interaction through Affordances

Authors: Mingyuan Lei, Yoonchang Sung, Tat-Jen Cham

Organizations: College of Computing and Data Science Nanyang Technological University, Singapore

Abstract

Generating realistic human-object interactions (HOI) in complex 3D scenes requires two complementary capabilities: reasoning about interaction feasibility in the environment and synthesizing realistic human-object motion. However, supervision for these capabilities is rarely available jointly at scale. Human-scene datasets provide rich information about environment-aware motion, while human-object datasets capture detailed interaction dynamics, yet paired human-object-scene data remain scarce. We present MAMHOI, an affordance-mediated factorization for scene-aware human-object interaction generation. MAMHOI factorizes scene-aware HOI generation through an explicit motion-affordance interface between scene understanding and motion synthesis: a scene-conditioned model first predicts where and how an interaction can be feasibly executed, and an affordance-conditioned HOI model then generates the corresponding human-object motion. This factorization allows scene understanding and interaction dynamics to be learned from complementary sources of supervision without requiring paired human-object-scene data. Experiments in complex indoor environments show that MAMHOI reduces object--scene penetration while better preserving human--object interaction quality, yielding more realistic and physically feasible scene-aware interactions. Project page: https://leimingyuan.github.io/MAMHOI-project-page/

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MaMi-HOI: Harmonizing Global Kinematics and Local Geometry for Human-Object Interaction Generation

    May 7, 2026Hao Wang, Shiqi Wang, Qi LiuHOI GenerationHuman Motion Generation

  2. Learning to Generate Human-Human-Object Interactions from Textual Descriptions

    Nov 25, 2025Jeonghyeon Na, Sangwon Baik, Inhee Lee +2HOI GenerationHuman Motion Generation

  3. PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph

    Aug 3, 2026Zhenhao Zhang, Jiajun Zhang, Wei Min +1HOI GenerationDexterous Manipulation