Harnessing Multimodal Large Language Models for Training-Free Human-Object Interaction Detection
Organizations: Shanghai Jiao Tong University · Xinjiang University · Beijing Jiaotong University · Shanghai AI Laboratory
Abstract
Human-object interaction (HOI) detection aims to localize human-object pairs and recognize their interactions. Traditional supervised methods perform strongly but rely on task-specific training. Recent multimodal large language models (MLLMs) offer a promising route to training-free HOI detection through their broad visual-semantic knowledge and versatile perceptual and reasoning capabilities. However, existing approaches largely invoke these capabilities through loosely coordinated inference stages. This fragmented execution restricts the role of interaction hypotheses in guiding visual exploration, leaving key participants overlooked and local ambiguities unresolved. Furthermore, propagating early semantic assumptions through subsequent visual grounding and relation prediction induces self-reinforcing semantic circularity. To resolve these challenges, we propose HarnessHOI, a training-free framework that transforms passive MLLM inference into an active interaction-centric harness. Specifically, we introduce an interaction-guided perception mechanism that projects emerging interaction hypotheses back into the visual space to discover missing participants and refine ambiguous evidence through targeted observation. Furthermore, a relation-agnostic geometric adjudication module reconciles multi-source evidence to establish a unified spatial basis for grounded interaction reasoning across multiple actions and semantic roles. Extensive experiments on HICO-DET and V-COCO demonstrate that HarnessHOI achieves state-of-the-art performance among training-free methods, confirming the effectiveness of the proposed harness for complex interaction understanding. Code will be released upon publication.
Figures & tables
| HICO-DET Default | HICO-DET Known Object | V-COCO | |||||||
| Method | Model | Full | Rare | Non-Rare | Full | Rare | Non-Rare | S1 | S2 |
| Task-optimized methods | |||||||||
| QPIC ( Tamura et al., 2021 ) | R101 | 29.90 | 23.92 | 31.69 | 32.38 | 26.06 | 34.27 | 58.30 | 60.70 |
| UPT ( Zhang et al., 2022a ) | R101-DC5 | 32.62 | 28.62 | 33.81 | 36.08 | 31.41 | 37.47 | 61.30 | 67.10 |
| GEN-VLKT ( Liao et al., 2022 ) | R101/CLIP-B | 34.95 | 31.18 | 36.08 | 38.22 | 34.36 | 39.37 | 63.58 | 65.93 |
| PViC ( Zhang et al., 2023 ) | Swin-L/H-DETR | 44.32 | 44.61 | 44.24 | 47.81 | 48.38 | 47.64 | 64.10 | 70.20 |
| RF-UC | NF-UC | UO | UV | |||||||||
| Method | Unseen | Seen | Full | Unseen | Seen | Full | Unseen | Seen | Full | Unseen | Seen | Full |
| Task-optimized methods | ||||||||||||
| GEN-VLKT ( Liao et al., 2022 ) | 21.36 | 32.91 | 30.56 | 25.05 | 23.38 | 23.71 | 10.51 | 28.92 | 25.63 | 20.96 | 30.23 | 28.74 |
| HOICLIP ( Ning et al., 2023 ) | 25.53 | 34.85 | 32.99 | 26.39 | 28.10 | 27.75 | 16.20 | 30.99 | 28.53 | 24.30 | 32.19 | 31.09 |
| CLIP4HOI ( Mao et al., 2023 ) | 28.47 | 35.48 | 34.08 | 31.44 | 28.26 | 28.90 | 31.79 | 32.73 | 32.58 | 26.02 | 31.14 | 30.42 |
| ADA-CM (FT) ( Lei et al., 2023 ) | 27.63 | 34.35 | 33.01 | 32.41 | 31.13 | 31.39 | – | – | – | – | – | – |
| Variant | Harness configuration | HICO-DET | V-COCO | ||||
| D | O | G | V | Full | S1 | S2 | |
| Direct HOI generation | – | – | – | – | 15.62 | 16.41 | 21.20 |
| Initial-instance baseline | – | – | – | 31.45 | 42.90 | 47.20 | |
| w/o discovered participants | 32.55 | 44.02 | 48.93 | ||||
| w/o targeted observation | – | 32.80 | 44.33 | 49.21 | |||
| First-candidate geometry | 32.95 | 44.38 | 49.48 | ||||
| Design factor | Alternative | HICO-DET | V-COCO | |
|---|---|---|---|---|
| Full | S1 | S2 | ||
| Participant discovery | Repeated localization | 32.45 | 43.87 | 48.88 |
| Local observation | Non-directed regions | 32.70 | 44.07 | 49.08 |
| Geometric adjudication | Relation-conditioned | 32.79 | 44.18 | 49.16 |
| Relation verification | Proposal-conditioned | 32.50 | 43.92 | 48.84 |
| HarnessHOI | 33.20 | 44.65 | 49.70 | |
Appendix figures & tables48 assets
Supplementary material from the paper’s appendix.
Appendix
| Operation | Visual and semantic context | MLLM operation | Shared outcome |
|---|---|---|---|
| Interaction-guided perception | |||
| Initial localization | Full image; object vocabulary | Generate participants in interaction groups | Initial entities and humans |
| Interaction discovery | Full image; initial entities; designated humans; public relations | Generate hypotheses, additional participants, and observation requests | Expanded entities and discovery associations |
| Targeted observation | One requested crop; target description; local instance context | Generate instance updates and newly visible participants | Candidate geometry and observation associations |
| Relation-agnostic geometric adjudication | |||
| Grounding selection | Full image and acquired crops; candidate boxes for one identity | Answer which candidate matches the visible instance | One grounding per identity |
| Step | Operation |
|---|---|
| 1 | Localize initial entities and identify the initial humans . Initialize each entity’s geometry candidates with its returned box. |
| 2 | Partition into groups of at most four. For each group, discover interactions and participants using the full image and . Associate the resulting hypotheses and observation requests with the expanded participant set. |
| 3 | Select at most two image-level observation requests. Localize participants in each requested crop, map their boxes to image coordinates, and update geometry candidates and observation associations. |
| 4 | For each multi-candidate identity, ask the geometry question over the full image and acquired crops. Select one grounding and construct the shared entity state . |
| 5 | Assess all directed human-entity pairs. Form from pair support, discovery associations, and observation associations. For HICO-DET, independently assess no-interaction for all pairs. |
| 6 | For HICO-DET, evaluate every compatible positive action on each retained pair through a visual question. For V-COCO, assess each human’s actions and role states, followed by role bindings for retained pairs. |
| Setting | Value |
|---|---|
| Model and visual inputs | |
| Backbone | Frozen Qwen3.5-9B; shared weights across stages |
| Precision | BF16 weights and KV cache |
| Maximum context | 16,384 tokens |
| Full-image processor | Roughly 3 MP minimum and 17 MP maximum, depending on aspect ratio |
| Crop processor | Roughly 0.8 MP |
| Default | Known Object | |||
|---|---|---|---|---|
| Method | Common | Original | Common | Original |
| Direct generation | 15.62 | 15.50 | 17.27 | 17.11 |
| Initial-instance baseline | 31.45 | 31.23 | 35.05 | 34.82 |
| w/o targeted observation | 32.80 | 32.56 | 36.39 | 36.15 |
| First-candidate geometry | 32.95 | 32.71 | 36.55 | 36.30 |
| HarnessHOI | 33.20 | 32.97 | 36.85 | 36.61 |
| S1 | S2 | |||
|---|---|---|---|---|
| Method | Common | Original | Common | Original |
| Direct generation | 16.41 | 16.41 | 21.20 | 21.20 |
| Joint role question | 41.10 | 41.10 | 46.30 | 46.30 |
| HarnessHOI | 44.65 | 44.65 | 49.70 | 49.70 |
| HICO-DET Default | HICO-DET Known Object | V-COCO | |||||||
| Method | Model | Full | Rare | Non-Rare | Full | Rare | Non-Rare | S1 | S2 |
| Task-optimized methods | |||||||||
| QPIC ( Tamura et al., 2021 ) | R101 | 29.90 | 23.92 | 31.69 | 32.38 | 26.06 | 34.27 | 58.30 | 60.70 |
| UPT ( Zhang et al., 2022a ) | R101-DC5 | 32.62 | 28.62 | 33.81 | 36.08 | 31.41 | 37.47 | 61.30 | 67.10 |
| GEN-VLKT ( Liao et al., 2022 ) | R101/CLIP-B | 34.95 | 31.18 | 36.08 | 38.22 | 34.36 | 39.37 | 63.58 | 65.93 |
| HOICLIP ( Ning et al., 2023 ) | R50/CLIP-B | 34.69 | 31.12 | 35.74 | 37.61 | 34.47 | 38.54 | 63.50 | 64.80 |
| RF-UC | NF-UC | UO | UV | |||||||||
| Method | Unseen | Seen | Full | Unseen | Seen | Full | Unseen | Seen | Full | Unseen | Seen | Full |
| Task-optimized methods | ||||||||||||
| GEN-VLKT ( Liao et al., 2022 ) | 21.36 | 32.91 | 30.56 | 25.05 | 23.38 | 23.71 | 10.51 | 28.92 | 25.63 | 20.96 | 30.23 | 28.74 |
| HOICLIP ( Ning et al., 2023 ) | 25.53 | 34.85 | 32.99 | 26.39 | 28.10 | 27.75 | 16.20 | 30.99 | 28.53 | 24.30 | 32.19 | 31.09 |
| CLIP4HOI ( Mao et al., 2023 ) | 28.47 | 35.48 | 34.08 | 31.44 | 28.26 | 28.90 | 31.79 | 32.73 | 32.58 | 26.02 | 31.14 | 30.42 |
| ADA-CM (FT) ( Lei et al., 2023 ) | 27.63 | 34.35 | 33.01 | 32.41 | 31.13 | 31.39 | – | – | – | – | – | – |
| System | Capability source | Task organization | Additional resources |
|---|---|---|---|
| HOI-R1 | SFT+GRPO-tuned MLLM | Joint interaction reasoning and localization | HOI instruction and reward supervision |
| DA-HOI (FT) | Adapted Qwen2.5-VL-3B | Spatial-aware pooling and deterministic matching | DETR or GroundingDINO; HOI supervision |
| DA-HOI (TF) | Frozen Qwen2.5-VL-3B | Candidate-conditioned deterministic generation | R50 DETR proposals |
| AgentHOI (Wang et al.) | BLIP and Qwen2.5 | Hypothesis proposal, multi-agent debate, and consensus | YOLOv8 |
| AgentHOI (Lei et al.) | GPT-4o or Qwen2.5-VL-72B | Recognition, re-mining, reassignment, grounding | GroundingDINO |
| DYSCO | CLIP/LongCLIP | Visual-semantic registry and scoring | R50 detector; target-domain visual support |
| HICO-DET metric | AgentHOI (Lei et al.) | HarnessHOI | Difference |
|---|---|---|---|
| Default Full | 29.68 | 33.20 | 3.52 |
| Default Rare | 40.49 | 41.63 | 1.14 |
| Default Non-Rare | 26.45 | 30.68 | 4.23 |
| Known Object Full | 33.71 | 36.85 | 3.14 |
| Known Object Rare | 43.39 | 44.20 | 0.81 |
| Known Object Non-Rare | 30.82 | 34.65 | 3.83 |
| RF-UC | NF-UC | UO | UV | |
|---|---|---|---|---|
| AgentHOI (Lei et al.), reported Unseen | 38.64 | 28.04 | 33.70 | 29.85 |
| AgentHOI (Lei et al.), common evaluator Unseen | 39.11 | 28.56 | 33.97 | 30.13 |
| HarnessHOI, Unseen | 40.15 | 33.64 | 35.45 | 33.87 |
| HarnessHOI, Seen | 31.46 | 33.09 | 32.75 | 33.09 |
| HarnessHOI, Full | 33.20 | 33.20 | 33.20 | 33.20 |
| Class group | Classes | GT | TP | Default | Known Object |
|---|---|---|---|---|---|
| Positive interactions | 520 | 29,110 | 19,965 | 36.84 | 39.79 |
| No interaction | 80 | 4,295 | 2,196 | 9.55 | 17.73 |
| Rare | 138 | 981 | 598 | 41.63 | 44.20 |
| Non-Rare | 462 | 32,424 | 21,563 | 30.68 | 34.65 |
| GT | TP | AP (%) | ||||
| Action / role | Physical | Missing | Physical | Missing | S1 | S2 |
| hold / obj | 2,733 | 875 | 2,234 | 854 | 42.95 | 52.39 |
| sit / instr | 1,089 | 827 | 747 | 786 | 38.93 | 47.88 |
| ride / instr | 505 | 51 | 428 | 47 | 63.13 | 63.40 |
| look / obj | 2,752 | 595 | 1,944 | 583 | 26.27 | 29.18 |
| hit / instr | 315 | 34 | 286 | 34 | 43.63 | 46.52 |
| Action | GT | TP | AP | Action | GT | TP | AP |
|---|---|---|---|---|---|---|---|
| hold | 3,608 | 3,533 | 85.42 | catch | 246 | 243 | 59.84 |
| stand | 4,118 | 4,013 | 78.45 | cut | 269 | 260 | 83.41 |
| sit | 1,916 | 1,838 | 84.87 | run | 687 | 681 | 80.02 |
| ride | 556 | 523 | 80.74 | work_on_computer | 410 | 393 | 80.02 |
| walk | 597 | 566 | 76.98 | ski | 424 | 420 | 88.57 |
| look | 3,347 | 3,276 | 43.27 | surf | 486 | 484 | 92.99 |
| Supervision | Backbone | HICO-DET Full mAP | V-COCO Role AP |
|---|---|---|---|
| Full | R50+Transformer | 34.10 | 65.44 |
| MLLM | R50+Transformer | 15.62 | 26.74 |
| MLLM | Swin-L+Transformer | 17.55 | 36.63 |
| mF1 | mPrec | mRec | |||
| Setting | Full | Rare | Non-Rare | Full | Full |
| HICO-DET | |||||
| End-to-end | 24.28 | 25.75 | 23.27 | 26.72 | 26.26 |
| R50-DETR | 30.32 | 27.12 | 31.68 | 35.71 | 36.21 |
| Oracle | 50.61 | 51.16 | 49.90 | 51.37 | 56.47 |
| V-COCO | |||||
| Annotation | Images | H-mIoU | O-mIoU | A-Acc | mAP |
|---|---|---|---|---|---|
| Fine-grained | Seen | ||||
| Unseen | |||||
| Precise | Seen | ||||
| Unseen | |||||
| Open-vocabulary | Seen | ||||
| Unseen |
| Comparison | Aspect examined | Common conditions |
|---|---|---|
| Overall inference | Direct generation, initial-instance baseline, and the complete harness | Backbone, full-image processing, public vocabulary, and evaluation |
| Perceptual components | Discovery participants, local observation, and geometric selection | Common initial localization; relation questions evaluated on each resulting participant set |
| Interaction guidance | Discovery cues, observation location, geometry criterion, or proposal context | Inputs and budgets specified for each replacement in Appendix D |
| Relation candidates | Pair sources and labels on the proposed pairs | Selected participants, image views, and relation supports |
| Role verification | Joint questions and compositional confidence | Physical and NULL candidates and participant geometry; joint questions use the corresponding binding or role-state views |
| Repeated generation | Response count and duplicate fusion | Ordered response prefixes and original generation scores |
| Default | Known Object | |||||
|---|---|---|---|---|---|---|
| Variant | Full | Rare | Non-Rare | Full | Rare | Non-Rare |
| HarnessHOI | 33.20 | 41.63 | 30.68 | 36.85 | 44.20 | 34.65 |
| Direct HOI generation | 15.62 | 22.13 | 13.68 | 17.27 | 23.25 | 15.49 |
| Initial-instance baseline | 31.45 | 39.79 | 28.96 | 35.05 | 41.30 | 33.18 |
| w/o discovered participants | 32.55 | 40.91 | 30.05 | 36.15 | 42.42 | 34.28 |
| w/o targeted observation | 32.80 | 41.24 | 30.28 | 36.39 | 42.62 | 34.53 |
| Variant | S1 | S2 |
|---|---|---|
| HarnessHOI | 44.65 | 49.70 |
| Direct HOI generation | 16.41 | 21.20 |
| Initial-instance baseline | 42.90 | 47.20 |
| w/o discovered participants | 44.02 | 48.93 |
| w/o targeted observation | 44.33 | 49.21 |
| First-candidate geometry | 44.38 | 49.48 |
| RF-UC | NF-UC | |||
|---|---|---|---|---|
| Variant | Unseen | Seen | Unseen | Seen |
| HarnessHOI | 40.15 | 31.46 | 33.64 | 33.09 |
| Direct generation | 21.82 | 14.07 | 12.80 | 16.33 |
| Initial-instance baseline | 38.33 | 29.73 | 31.59 | 31.42 |
| w/o discovered participants | 39.17 | 30.89 | 33.26 | 32.37 |
| w/o targeted observation | 40.00 | 31.00 | 32.88 | 32.78 |
| Participant and relation domain | HICO-DET | V-COCO |
|---|---|---|
| Initial participant groundings | 22,314 | 14,248 |
| After interaction discovery | 22,885 | 14,465 |
| All localization candidates | 22,897 | 14,488 |
| Selected shared groundings | 22,895 | 14,473 |
| Final action or role candidates | 22,232 | 13,628 |
| Annotated interactions or roles | 33,405 | 16,392 |
| Interaction group | Annotations | Initial-instance baseline | HarnessHOI |
|---|---|---|---|
| HICO-DET | |||
| All interactions | 33,405 | 21,354 | 22,232 |
| Multiple annotated human boxes | 30,441 | 19,498 | 20,288 |
| Multiple targets for the same action | 11,549 | 6,172 | 6,437 |
| Small targets | 5,960 | 2,184 | 2,356 |
| V-COCO | |||
| Group | GT | Initial | Discovery | Observed | Selected | Relations |
|---|---|---|---|---|---|---|
| HICO-DET | ||||||
| All interactions | 33,405 | 22,314 | 22,885 | 22,897 | 22,895 | 22,232 |
| Multiple annotated human boxes | 30,441 | 20,323 | 20,848 | 20,858 | 20,856 | 20,288 |
| Multiple targets per action | 11,549 | 6,578 | 6,728 | 6,733 | 6,731 | 6,437 |
| Small targets | 5,960 | 2,387 | 2,489 | 2,493 | 2,492 | 2,356 |
| V-COCO | ||||||
| HICO-DET | V-COCO | ||
| Variant | Full | S1 | S2 |
| (a) Sources of relation candidates | |||
| Discovery and observation | 28.09 | 42.48 | 47.19 |
| Visual pair assessment | 32.30 | 43.23 | 47.77 |
| Combined | 33.20 | 44.65 | 49.70 |
| (b) Labels retained on identical proposed pairs | |||
| S1 | S2 | |||||
|---|---|---|---|---|---|---|
| Action / role | Joint | Compositional | Joint | Compositional | ||
| hold / obj | 46.08 | 42.95 | -3.13 | 52.08 | 52.39 | +0.31 |
| sit / instr | 38.97 | 38.93 | -0.04 | 45.02 | 47.88 | +2.86 |
| ride / instr | 56.50 | 63.13 | +6.63 | 57.05 | 63.40 | +6.35 |
| look / obj | 31.49 | 26.27 | -5.22 | 35.52 | 29.18 | -6.34 |
| hit / instr | 27.68 | 43.63 | +15.95 | 30.64 | 46.52 | +15.88 |
| Backbone | HICO-DET Full | V-COCO S1 | V-COCO S2 |
|---|---|---|---|
| Qwen3.5-9B | 33.20 | 44.65 | 49.70 |
| InternVL3-8B | 32.11 | 43.60 | 48.40 |
| Qwen2.5-VL-7B | 31.61 | 43.10 | 47.90 |
| Quantity | HICO-DET | V-COCO |
|---|---|---|
| Perceptual generation | ||
| Generative responses (mean) | 2.51 | 2.78 |
| Generative responses (median) | 2 | 2 |
| Generative responses (P95) | 5 | 6 |
| Output tokens (mean) | 563.29 | 615.26 |
| Local observations (mean) | 0.19 | 0.40 |
| Scoring task | HICO-DET | V-COCO |
|---|---|---|
| Geometry selection | 0.17 | 0.33 |
| Pair assessment | 43.10 | 50.82 |
| No-interaction assessment | 43.10 | – |
| Positive-action verification | 33.44 | – |
| Agent assessment | – | 86.73 |
| Role visibility | – | 83.40 |
| HICO-DET | V-COCO | ||||
|---|---|---|---|---|---|
| Method | Full | Cost | S1 | S2 | Cost |
| Direct generation, | 15.61 | 2.50 | 16.38 | 21.18 | 2.60 |
| Direct generation, | 17.62 | 4.98 | 19.39 | 24.19 | 5.22 |
| Direct generation, | 19.04 | 9.97 | 20.86 | 25.54 | 10.50 |
| Direct generation, | 19.84 | 19.95 | 22.21 | 26.50 | 20.94 |
| HarnessHOI | 33.20 | 13.26 | 44.65 | 49.70 | 22.07 |
| Fusion | Responses | HICO-DET Full | V-COCO S1 | V-COCO S2 |
|---|---|---|---|---|
| Exact-match fusion | 1 | 15.62 | 16.41 | 21.20 |
| 2 | 14.53 | 13.31 | 16.55 | |
| 4 | 12.78 | 9.43 | 11.61 | |
| 8 | 10.87 | 6.51 | 7.84 | |
| Pair-IoU suppression | 1 | 15.61 | 16.38 | 21.18 |
| 2 | 17.62 | 19.39 | 24.19 |
| Method | HICO-DET Full | V-COCO S1 | V-COCO S2 |
|---|---|---|---|
| Direct HOI generation | 15.62 | 16.41 | 21.20 |
| Direct HOI generation, GT ranking | 21.58 | 30.69 | 35.01 |
| HarnessHOI | 33.20 | 44.65 | 49.70 |
| Candidate construction | Responses | HICO-DET Full | V-COCO S1 | V-COCO S2 |
|---|---|---|---|---|
| Exact-match fusion | 1 | 21.58 | 30.69 | 35.01 |
| 2 | 24.30 | 34.80 | 38.90 | |
| 4 | 27.80 | 39.20 | 43.10 | |
| 8 | 30.50 | 41.90 | 46.40 | |
| Pair-IoU suppression | 1 | 21.55 | 30.58 | 34.90 |
| 2 | 24.09 | 34.50 | 38.60 |
| Method | Seed | HICO-DET Full | V-COCO S1 | V-COCO S2 |
|---|---|---|---|---|
| Direct generation | 66 | 15.62 | 16.41 | 21.20 |
| 67 | 15.81 | 16.89 | 21.93 | |
| 68 | 15.55 | 16.58 | 21.74 | |
| Mean SD | 15.66 0.13 | 16.63 0.24 | 21.62 0.38 | |
| HarnessHOI | 66 | 33.20 | 44.65 | 49.70 |
| 67 | 32.79 | 44.80 | 49.75 |
| Default | Known Object | |||||
| Seed | Full | Rare | Non-Rare | Full | Rare | Non-Rare |
| 66 | 33.20 | 41.63 | 30.68 | 36.85 | 44.20 | 34.65 |
| 67 | 32.79 | 41.04 | 30.32 | 36.39 | 43.74 | 34.20 |
| 68 | 32.86 | 41.17 | 30.38 | 36.50 | 43.87 | 34.30 |
| Mean | 32.95 | 41.28 | 30.46 | 36.58 | 43.94 | 34.38 |
| SD | 0.22 | 0.31 | 0.19 | 0.24 | 0.24 | 0.24 |
| Case | Action and target role | Person IoU | Target IoU | Score |
|---|---|---|---|---|
| Book | hold book | 0.9419 | 0.6543 | 0.9859 |
| read book | 0.9545 | 0.7631 | 0.9796 | |
| Shared bench | sit_on bench | 0.9497 | 0.8209 | 0.7959 |
| sit_on bench | 0.8886 | 0.8209 | 0.8658 | |
| Sandwich | hold / obj | 0.9659 | 0.8844 | 0.8896 |
| eat / obj | 0.9659 | 0.8844 | 0.8616 |
| Question | False / missing | True / visible | Unknown |
|---|---|---|---|
| Observer performs look | 0.2644 | 0.7187 | 0.0169 |
| Visibility of look/obj | 0.4069 | 0.5224 | 0.0707 |
| Observer-target pair | 0.7546 | 0.2162 | 0.0293 |
| Case | Action / role | ||||
|---|---|---|---|---|---|
| Sandwich | hold / obj | 0.9770 | 0.9769 | 0.9320 | 0.8896 |
| eat / obj | 0.9890 | 0.9525 | 0.9146 | 0.8616 | |
| Tennis | hold / obj | 0.9399 | 0.9669 | 0.9627 | 0.8748 |
| hit / instr | 0.9740 | 0.9876 | 0.9241 | 0.8889 | |
| hit / obj | 0.9740 | 0.9914 | 0.9241 | 0.8924 | |
| Cutting | hold / obj | 0.9323 | 0.9234 | 0.9241 | 0.7955 |
| Question | Visual judgment | Answer set |
|---|---|---|
| Pair | Whether any legal positive action or physical-role binding connects the specified human and target, including distant and non-contact relations | false / true / unknown |
| No interaction | Whether both instances and their boxes are valid and no legal positive interaction holds between them | false / true / unknown |
| Action | Whether the specified participant pair and action jointly have visual support | false / true / unknown |
| Agent | Whether the specified human performs the action, independently of target visibility | false / true / unknown |
| Role state | Given the action, whether its role has a localizable visible target, has a missing target, or remains unresolved | missing / visible / unknown |
| Binding | Given the action and visible role, whether this particular target fulfills that role | false / true / unknown |
| Action | Role | Target meaning |
|---|---|---|
| hold | obj | Object held by the person |
| sit | instr | Seat or surface supporting the sitting person |
| ride | instr | Vehicle or animal being ridden |
| look | obj | Target of the person’s gaze |
| hit | instr | Implement used for hitting, such as a racket or bat |
| hit | obj | Target being hit, such as a ball |