When humans hand each other objects, they incorporate both geometric and semantic information into this process. For example, passing a knife with the handle towards the recipient, rather than the blade, is both more ergonomic and safer. Recent state-of-the-art methods for task-oriented robot-human handovers have progressed from modeling object geometry to incorporating object affordances. However, they often forgo predicting the explicit, task-specific hand poses a human selects to utilize an object. Since many objects support multiple interaction modalities, e.g., a claw hammer used to strike or pull nails, this variability must be modeled to achieve robust task-oriented handovers. To tackle this, we propose a novel approach, GENESIS-Handover (GENErative HypotheSIS), which leverages VLM image generation to produce a variety of task-specific hand-object interaction hypotheses. These hypotheses are matched in real time to the observed human hand pose, enabling inference of the most suitable handover configuration. By leveraging VLMs as priors of plausible hand-object interactions, the method produces task-conditioned handover strategies for previously unseen object-task pairs. We evaluate the standalone interaction proposal module before deploying the full system on a mobile manipulator. In a user study with 12 participants across five task-object pairs, 83.3% perceived our method to have better task understanding than the previous state of the art.
Figures & tables
Figure 1: Example for the tool-task pair screwdriver, to screw . Initially, GENESIS-Handover generates multiple hand-object interaction hypotheses via a VLM. At handover onset (I), the human hand pose is observed by the onboard camera, and each hypothesis is scored based on joint-level similarity to the observed pose. The closest hypothesis is selected for execution. During execution, the robot continuously adapts its target to the human hand (II and III, hand pose from previous frame blended in). Finally, the handover occurs with the hand configuration matching the selected hypothesis (IV).
Figure 2: Overview of GENESIS-Handover. The system generates task-conditioned hand-object interaction hypotheses (blue), selects the best-matching hypothesis based on the observed human hand (purple), and computes a reactive end-effector target by anchoring the selected relation to the live hand pose (dark green). AFT-Handover provides affordance priors for alignment initialization and object placement in the gripper (white). All inputs to the system are shown in light green.
Figure 3: Overview of the Hand-Object Hypothesis Generation block. From an RGB-D scene observation and task specification, the system generates multiple task-conditioned hand-object hypotheses and maps them into the real object frame: The system first reconstructs the object mesh (I), then generates multiple synthetic hand-object interaction hypothesis images (II), generates depth and hand pose estimates (III and IV), lifts these into metric 3D (V), and expresses each hypothesis as a relative transform between the human hand and the real object (VI).
Metric
Mug
Pan
Apple
Hammer
Knife
Cellphone
Cup
Flashlight
Scissors
Toothbrush
Teapot
Wine Glass
H. Div.
15.1
10.5
28.0
10.4
10.2
9.5
31.1
8.7
18.3
12.9
12.1
40.0
Gen → CP ↓
8.3
11.6
14.3
8.7
8.8
9.1
12.8
11.4
9.0
11.5
11.8
9.1
CP → Gen ↓
11.3
12.5
16.2
10.2
9.2
9.2
14.7
11.8
10.7
10.3
10.6
14.5
Table 1: Grasp distribution analysis. Distances are symmetric surface distances (in mm).
Method
Perfect
Small
Big
Fail
GENESIS-Handover
16
30
11
3
AFT-Handover
6
28
21
5
Table 2: Handover distribution for GENESIS-Handover and AFT-Handover for the 60 user study handovers each.
Figure 4: Example images comparing hand alignment between AFT-Handover and the proposed method for all handovers of one user study participant. GENESIS-Handover more closely matches the human hand configuration, requiring fewer adjustments to grasp the object than AFT-Handover. The initial handover pose is shown in transparent.
Object
n
Δ Rand. ↑
Oracle Gap ↑
Better Rand. ↑
Oracle Match ↑
Hammer
8
16.6%
77.8%
100.0%
75.0%
Ladle
9
12.4%
62.0%
75.0%
62.5%
Mug
7
19.3%
35.3%
80.0%
0.0%
Paintbrush
7
14.9%
54.7%
85.7%
42.9%
Screwdriver
6
19.2%
49.2%
75.0%
50.0%
Overall
37
16.2%
56.9%
83.3%
47.5%
Table 3: Per-object performance of hypothesis selection. Δ Rand. denotes the relative reduction in average per-joint error compared to the mean random baseline. Oracle Gap indicates the fraction of the random-to-oracle gap recovered. Better Rand. reports the percentage of cases where the selected hypothesis achieves a lower error than the random baseline. Oracle Match indicates the percentage of cases where the selected hypothesis coincides with the oracle.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Object
H. Var.
Gen → CP ↓
CP → Gen ↓
Dgen↑
H. Div.
Mug
6.4
8.3
11.3
18.6
15.1
Pan
6.0
11.6
12.5
15.8
10.5
Apple
10.6
14.3
16.2
26.1
28.0
Hammer
6.0
8.7
10.2
11.9
10.4
Knife
6.0
8.8
9.2
15.4
10.2
Bottle
15.1
12.4
13.5
30.8
24.3
Appendix
Table A.1: Grasp distribution analysis comparing generated grasps to ContactPose. Distances are symmetric surface distances (in mm).
GENESIS-Handover (aggregate)
AFT-Handover (aggregate)
Object
Prec. ↑
Rec. ↑
Corr. ↑
Area ↓
Prec. ↑
Rec. ↑
Corr. ↑
Area ↓
Mug
0.96
0.08
0.62
0.05
0.89
0.14
0.59
0.10
Hammer
1.00
0.23
0.57
0.22
1.00
0.32
0.91
0.31
Pan
1.00
0.04
0.84
0.04
1.00
0.07
0.98
0.08
Knife
1.00
0.40
0.81
0.39
1.00
0.38
0.86
0.38
Bottle
1.00
0.40
0.44
0.33
0.98
0.47
0.67
0.40
Appendix
Table A.2: Comparison to AFT-Handover [ 2 ] on aggregate contact prediction. We report precision, recall, Pearson correlation with ContactPose, and active contact area.
Figure A.1: Generated image examples. 1 (left head instead of right) and 4 (infeasible grasp) are corrected by WiLoR, as the right is matched correctly. Images 2 and 3 lead to failures as the grasps are wrong.
Figure A.2: Mean Likert-scale rating (with 1 being strongly disagree and 5 strongly agree) for different object-task pairs and evaluation criteria.
Object
Hand Config.
Hand Config. + RGB
Mean
Std
Min
Max
Mean
Std
Min
Max
Hammer
1.35
0.49
1
2
2.81
0.54
2
4
Ladle
1.67
0.62
1
3
2.80
0.41
2
3
Mug
2.63
0.50
2
3
2.19
0.40
2
4
Paintbrush
1.13
0.35
1
2
2.63
0.59
2
3
Screwdriver
1.94
0.25
2
2
2.88
0.62
2
4
Appendix
Table A.3: Hypothesis diversity measured as the number of clusters per handover. Clusters are obtained via hierarchical clustering (average linkage) with a fixed distance threshold. Hand clusters are computed from hand-configuration distances, while combined clusters additionally incorporate normalized RGB hand-object relative position. Values report mean cluster counts over handovers.
Object
Ambig. (%)
#Hyp.
#Clusters
Intra (%)
Hammer
35.3
1.88
1.06
94.0
Ladle
46.7
1.87
1.00
100.0
Mug
18.8
1.19
1.00
100.0
Paintbrush
81.3
2.25
1.00
100.0
Screwdriver
56.3
2.00
1.06
94.0
Average
47.5
1.84
1.02
98.0
Appendix
Table A.4: Ambiguity statistics per object. We report the fraction of ambiguous handovers, the average number of competing hypotheses within the ambiguity threshold, and whether ambiguity occurs within a single cluster or across clusters.
Figure A.3: Failure case example on a mug. Among multiple plausible grasp hypotheses, the oracle grasp (green) aligns with the human’s final grasp, while the selected hypothesis (blue) corresponds to another strategy. This highlights the difficulty of ranking valid grasps from initial observations under ambiguity.
Figure A.4: Near-miss case for a mug. The selected hypothesis (blue) closely matches the oracle (green) and the final human grasp, with only minor differences. The initial hand pose already reflects the final grasp configuration, reducing ambiguity and enabling accurate selection.
Human-to-robot (H2R) object handover is a fundamental capability for human-robot collaboration, yet progress is hindered by the scarcity of large-scale, human-centric datasets and the significant sim-to-real gap. To address these challenges, we introduce Hand2Bot, an RGB-D video dataset that provides rich contextual information such as body posture and facial expressions, specifically collected for handover scenarios with real-world noise patterns. We further propose PassGen, a generative pipeline that leverages stable video diffusion and an Intention-Aware Temporal Face Encoder to synthesize realistic handover sequences while ensuring hand-object consistency. To bridge the sim-to-real gap, we implement a morphology-based depth editing strategy that replicates realistic sensor noise found in physical depth maps. Experimental evaluations demonstrate that our framework achieves high intention identification accuracy and low false trigger rates in both ablation studies and real-world deployment on a physical robot platform. Our results confirm that training on PassGen allows for robust zero-shot transfer and earlier intention anticipation compared to traditional hand-centric baselines, effectively enabling socially aware robotic behavior in shared workspaces.
Tianyu Sun, Zhoujie Fu, Zihui Gao +2
College of Computing and Data Science, Nanyang Technological University · College of Computer Science and Technology, Zhejiang University · Alibaba Group
Robot-to-human handovers often rely on static, open-loop strategies (or, at best, approaches that adapt only the position), which generally do not consider how the object will be grasped by the human, thus requiring the user to adapt. This work presents a novel adaptive framework that dynamically adjusts the object's delivery pose in real time based on the user's hand pose and the intended downstream task. By integrating AI-based hand pose estimation with smooth, kinematically constrained trajectories, the system ensures a safe approach and an optimal handover orientation. A comprehensive user study compares the proposed adaptive approach against a static baseline across multiple tasks, evaluating both subjective metrics (NASA-TLX, Human-Robot Trust Scale) and objective physiological data (blink rate measured via wearable eye-trackers). The results demonstrate that dynamic alignment significantly reduces users' cognitive workload and physiological stress, while improving their confidence in the robot's reliability. These findings highlight the potential of task- and pose-aware systems for enabling fluid and ergonomic human-robot collaboration.
Federico Biagi, Dario Onfiani, Simone Silenzi +2
Department of Engineering "Enzo Ferrari", University of Modena and Reggio Emilia, Italy. · Department of Surgery, Medicine, Dentistry and Morphological Sciences, University of Modena and Reggio Emilia, Italy.
We present R2HandoverSim, a simulation benchmark for robot-to-human (R2H) object handovers. Although R2H handover methods have advanced rapidly, the lack of standardized evaluation protocols impedes objective comparison. Our benchmark enables reproducible evaluation by systematically comparing four baselines on their predicted shared grasp poses. We conduct a user study with 30 participants, analyze baseline performance, and show that simulation results correlate with real-world evaluation outcomes. Crucially, five complementary metrics (planning feasibility, reachability, grasp stability, grasp affordance, and safety) better reflect user-perceived handover quality than overall success rate alone. Website and code: https://robot-future.github.io/r2handoversim/.
Hanxin Zhang, Abdulqader Dhafer, Hongbiao Dong +1
DANiLab, University of Leicester, Leicester, UK · School of Computing and Mathematical Sciences, University of Leicester, Leicester, UK · School of Metallurgy and Materials, University of Birmingham, Birmingham, UK