cs.CVSep 28, 2026

HEIR: Learning Human-Entity Interactions with Functional Roles

Authors: Di Wen, Wenhao Guo, Yuedong Tan, Yun Huang, Minheng Wu, Zhihang Chen, Haiwen Sun, Fei Teng, +10 more

Organizations: Karlsruhe Institute of Technology (KIT) · Institute for Computer Science, Artificial Intelligence and Technology (INSAIT) · Hunan University · University of Bremen · ETH Zurich

Abstract

Understanding human-entity interactions requires recovering each person-action event's participants, roles, and shared identities. This structure can support embodied agents by clarifying who acts on which entities and how, informing anticipation and coordination in shared environments. Standard HOI metrics score individual links, leaving complete event composition undermeasured. We introduce HEIR (Human-Entity Interactions with Functional Roles), an image benchmark for complete grounded participant-role sets across object, interpersonal, and self-directed interactions. It contains 18,730 images, six roles, 105 actions, and 437 nouns, with shared entities, role changes, and repeated fillers; 51.6% of images contain multiple actors and 62.1% contain multiple actions. HEIR pairs relation AP with complete-set AP and structural evaluation. We also introduce CoRISP (Compositional Role-aware Interaction Set Prediction), which uses shared entity identities to combine role-conditioned evidence and predict normalized participant-role sets. Cardinality and role-multiplicity potentials couple assignments through event size and role composition, with exact per-event normalization. Across 16 baselines, relation and complete-event rankings diverge even after aligning action weights. CoRISP leads the evaluated systems on repeated-role events and shared-participant images in HEIR by 2.87 and 3.82 Set mAP points, respectively. On V-COCO, CoRISP achieves 73.72/76.23 role AP and 61.06/68.59 complete-set AP on two-slot actions under Scenarios 1/2. These results show the value of learning and evaluating event composition alongside individual relations. The code and dataset are publicly available at https://github.com/Kratos-Wen/HEIR.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation

    Jul 5, 2026Yaozu Wu, Wei-Chieh Huang, Jizhou Guo +11Collaboration

  2. IAE-VTG: Interaction-Aligned Action-Entity Video Temporal Grounding

    Sep 9, 2026Shiwen Zhao, Qi Zhang, Sezer Karaoglu +2Video Temporal GroundingVision-Language Model Grounding

  3. Learning to Generate Human-Human-Object Interactions from Textual Descriptions

    Nov 25, 2025Jeonghyeon Na, Sangwon Baik, Inhee Lee +2Articulated ObjectsHuman Motion Generation