HEIR: Learning Human-Entity Interactions with Functional Roles
Authors: Di Wen, Wenhao Guo, Yuedong Tan, Yun Huang, Minheng Wu, Zhihang Chen, Haiwen Sun, Fei Teng, +10 more
Organizations: Karlsruhe Institute of Technology (KIT) · Institute for Computer Science, Artificial Intelligence and Technology (INSAIT) · Hunan University · University of Bremen · ETH Zurich
Understanding human-entity interactions requires recovering each person-action event's participants, roles, and shared identities. This structure can support embodied agents by clarifying who acts on which entities and how, informing anticipation and coordination in shared environments. Standard HOI metrics score individual links, leaving complete event composition undermeasured. We introduce HEIR (Human-Entity Interactions with Functional Roles), an image benchmark for complete grounded participant-role sets across object, interpersonal, and self-directed interactions. It contains 18,730 images, six roles, 105 actions, and 437 nouns, with shared entities, role changes, and repeated fillers; 51.6% of images contain multiple actors and 62.1% contain multiple actions. HEIR pairs relation AP with complete-set AP and structural evaluation. We also introduce CoRISP (Compositional Role-aware Interaction Set Prediction), which uses shared entity identities to combine role-conditioned evidence and predict normalized participant-role sets. Cardinality and role-multiplicity potentials couple assignments through event size and role composition, with exact per-event normalization. Across 16 baselines, relation and complete-event rankings diverge even after aligning action weights. CoRISP leads the evaluated systems on repeated-role events and shared-participant images in HEIR by 2.87 and 3.82 Set mAP points, respectively. On V-COCO, CoRISP achieves 73.72/76.23 role AP and 61.06/68.59 complete-set AP on two-slot actions under Scenarios 1/2. These results show the value of learning and evaluating event composition alongside individual relations. The code and dataset are publicly available at https://github.com/Kratos-Wen/HEIR.
Figures & tables
Figure 1: HEIR structure. (a) Shared participants across events. (b) Interaction structure (top): HHI denotes images with at least one human–human relation; other images are grouped by acting-person count versus distinct participant count. Bottom: actor counts. All split images are included. (c) Role shares for the 24 most frequent actions. (d) Training support of 2,972 observed classes, including 145 unseen classes with zero training relations.
Figure 2: CoRISP overview. (a) Role-conditioned recurrent updates. (b) Role-preserving aggregation over event, pair, and shared-entity neighborhoods. (c) Exact set normalization with cardinality and role-multiplicity potentials. Marginals score individual relations; joint probabilities score complete assignments. The displayed assignments illustrate the output structure.
HEIR
V-COCO (S1/S2)
Method
Params (M) ↓
HOI ↑
Role ↑
Set mAP ↑
AProle↑
Set mAP ↑
Full
Single
Multi
Repeat
Shared
All
Dual
Visual models
QPIC R50 ( Tamura et al., 2021 ) †
41.6
15.11
14.82
14.14
16.03
4.92
3.68
14.71
58.79/60.97 W
53.87/57.77
41.64/45.84
QPIC R101 ( Tamura et al., 2021 ) †
60.5
15.57
15.26
14.84
16.93
3.86
3.01
15.10
58.16/60.65 W
53.03/57.11
39.38/44.28
MUREN ( Kim et al., 2023 ) †
75.1
16.38
16.17
13.98
16.04
6.57
4.66
15.75
68.72/70.97 W
62.97/67.33
54.94/60.95
Table 1: HEIR (left) and V-COCO (right) (%). V-COCO pairs are S1/S2; re-evaluated role AP excludes point . Params: HEIR trainable millions. † : shared HEIR inventory; W/T/A/R: released weights, trained official code, adapted official code, or reimplementation. Bold/underline: best/second, excluding literature-only scores.
Figure 3: Role ambiguity and event recovery across all 16 baselines. (a) HOI–Role AP gaps with equal action–noun pair weights. (b) Role and Set AP with equal weights over the same 101 actions; sets use common-prefix decoding.
Variant
Role Full ↑
Set mAP ↑
Full
Single
Multi
Repeat
Shared
w/o G
20.67
17.09
20.54
6.89
5.39
19.96
w/o context messages
20.93
16.42
19.40
6.70
4.38
19.75
w/o role feedback
20.95
17.05
20.18
6.97
4.17
19.97
CoRISP
20.41
18.04
20.57
8.94
8.99
23.80
Table 2: HEIR component ablations (%). The model and evaluation settings follow Table 1
Variant
Role Full ↑
Set mAP ↑
Full
Single
Multi
Repeat
Shared
w/o G
20.67
17.09
20.54
6.89
5.39
19.96
w/o context messages
20.93
16.42
19.40
6.70
4.38
19.75
w/o role feedback
20.95
17.05
20.18
6.97
4.17
19.97
CoRISP
20.41
18.04
20.57
8.94
8.99
23.80
Table 2: HEIR component ablations (%). The model and evaluation settings follow Table 1
Method
Multi
Repeat
C↑
R↑
C↑
R↑
SOV-STG Swin-L ( Chen et al., 2025 )
51.21
24.22
56.06
24.94
PViC Swin-L ( Zhang et al., 2023 )
47.47
22.55
60.41
26.89
RLIPv2 Swin-L ( Yuan et al., 2023 )
48.30
23.94
54.58
25.51
CoRISP
48.92
27.48
62.01
33.41
Table 3: Action–cardinality oracle (%). The true action and count are supplied. C : coverage; R : recovery. All use 100 relations/image; emphasis ranks shown models.
Figure 4: Composition in four-person scenes. Highest-scoring set for the indicated actor/action, on identical uncropped images. Gray IDs mark reference participants; solid yellow/blue boxes are predicted actors/targets. In (a), CoRISP recovers three targets sharing one role; in (b), all models omit two participants in an occluded group hug.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Source
Released
What it contributes
Everyday
HICO-DET ( Chao et al., 2018 )
6,347
Web photographs, 117 actions, COCO nouns
V-COCO ( Gupta and Malik, 2015 )
1,848
COCO images with action-specific roles
SWiG ( Pratt et al., 2020 )
1,828
Situations with grounded semantic roles
Visual Genome ( Krishna et al., 2017 )
216
Dense scene-graph relations
Open Images ( Kuznetsova et al., 2020 )
191
Visual relationship annotations
Wikimedia Commons ( Wikimedia Foundation, 2026 )
157
Openly licensed photographs
Appendix
Table 4: Source datasets and collections. Released images per source and the setting each source contributes. HICO-DET, V-COCO, NVI, and SWiG enter with native interaction labels; all other sources are annotated from model proposals.
Readout
Full ↑
Single best nonempty assignment
17.56
State-wise MAP, K=8
18.04
Appendix
Table 5: Effect of the set hypothesis budget (Set mAP, %). The CoRISP checkpoint is fixed. K counts state winners; both settings retain at most 100 sets/image.
Method
Multi
Repeat
C↑
R↑
C↑
R↑
Visual models
QPIC R50 ( Tamura et al., 2021 )
28.94
14.99
31.69
16.25
QPIC R101 ( Tamura et al., 2021 )
31.58
15.75
33.98
16.48
MUREN ( Kim et al., 2023 )
34.91
16.79
40.16
17.85
SOV-STG-S ( Chen et al., 2025 )
42.19
17.49
49.08
18.88
Appendix
Table 6: Action–cardinality oracle at 100 relations per image (%). C : box-and-noun coverage; R : exact top- k⋆ recovery with the true action and count supplied. Denominators include all 1,441 Multi and 874 Repeat events. Denominators include all 1,441 Multi and 874 Repeat events.
Method
Encoder(s)
Reported role AP ↑
Visual models
QPIC ( Tamura et al., 2021 )
R50
58.8/61.0
QPIC ( Tamura et al., 2021 )
R101
58.3/60.7
MUREN ( Kim et al., 2023 )
R50
68.8/71.0
SOV-STG-S ( Chen et al., 2025 )
R50
–
SOV-STG R101 ( Chen et al., 2025 )
R101
63.9/65.4
Appendix
Table 7: Published V-COCO role AP (%, S1/S2). Source averaging conventions differ; scores are therefore unranked. R/C/D: ResNet/CLIP/DINOv3; DDETR: Deformable DETR.
School of Computer Science, The University of Sydney, Sydney, NSW 2006, Australia · Informatics Institute at University of Amsterdam, 1098 XH Amsterdam, the Netherlands