cs.CVSep 21, 2026

Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes

Authors: Hanyang Kong, Xingyi Yang

Organizations: National University of Singapore · The Hong Kong Polytechnic University

Abstract

Understanding interaction in a 3D scene requires recovering movable parts, their motion, and where they can be operated. These quantities are related, and their predictions can inform one another. A closed cabinet door, for instance, reveals a movable surface but may leave the hinge side ambiguous; its handle helps resolve this ambiguity, while the part provides context for localizing and interpreting the small handle. Building on this observation, we present SEGMENT-SNAP, which combines geometric and semantic evidence through part-handle coupling. Three independently trained predictors recover movable parts, dense handles, and part-associated handle proposals. We couple their outputs in two directions. For part motion, a training-free geometric decoder fits predicted part surfaces under explicit physical priors and uses detected handles to select candidate hinge lines. For handle prediction, a part-conditioned branch proposes additional handles, while standalone part classes refine their rotation/translation labels, with dense-handle labels as a fallback. Each transfer is applied once, without iterative feedback. On the Articulate3D validation set, handle guidance raises motion-gated AP from 13.74 to 40.98 under fixed masks and axes. Additional handle proposals raise handle AP from 24.63 to 29.65, and full contextual class correction raises it to 30.99 in the reference configuration. Fixed-input controls, retraining ablations, learned-decoder comparisons, and paired visualizations together characterize the benefits and limits of this coupling. Our system also achieved first place in the Articulate3D Challenge.

Figures & tables

Appendix figures & tables20 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph

    Aug 3, 2026Zhenhao Zhang, Jiajun Zhang, Wei Min +1Hand-Object Interaction Detection3D Hand

  2. GRAFT: Geometric Refinement and Fitting Transformer for Human Scene Reconstruction

    Apr 21, 2026Pradyumna YM, Yuxuan Xue, Yue Chen +3Scene ReconstructionArticulated Objects

  3. DeSeG: Decoupling Semantic Intent and Geometric Constraints for Physically Plausible Human-Scene Interaction

    Jul 7, 2026Jiakun Li, Zhe Li, Wenqiang Wu +4Articulated ObjectsGeometric Constraints