Function beyond Form: Functional Correspondence for Cross-Embodiment Dexterous Grasp Generation
Authors: Bolin Zou, Wenlong Dong, Mu Ai, Chao Tang, Aoxiang Gu, Lipeng Chen, Hong Zhang
Organizations: Shenzhen Key Laboratory of Robotics and Computer Vision, Southern University of Science and Technology, Shenzhen, China. · Department of Robotics, Perception and Learning, KTH Royal Institute of Technology, Stockholm, Sweden. · School of Artificial Intelligence, Shanghai Jiao Tong University, Shanghai, China. · Rysen Robotics, Shenzhen, China.
Cross-embodiment dexterous grasp generation remains challenging because robotic hands differ substantially in geometry, topology, and kinematics. Existing approaches often lack explicit correspondences between structurally different hand regions that play similar functional roles in a grasp, a concept we refer to as functional correspondence. Consequently, their models tend to learn hand-specific interaction patterns rather than transferable grasp knowledge, limiting generalization to unseen hands. To address this limitation, we introduce FunCo-Grasp, which establishes functional correspondences across heterogeneous hand embodiments. Specifically, Functional Part Alignment aligns each hand to a canonical functional schema by mapping physical links to shared functional parts according to their grasping roles, while Canonical Frame Alignment expresses these parts in canonical local frames. These two alignments provide a consistent representation for inter-part and hand-object interactions, allowing the model to learn transferable grasp knowledge across hands. Conditioned on the aligned hand representation and object geometry, a diffusion model generates the target spatial arrangement of the functional parts, which are then converted into an executable joint configuration. Adapting FunCo-Grasp to an unseen hand requires only its geometric and kinematic models and a one-time lightweight functional annotation, without target-hand grasp data, fine-tuning, or learned retargeting. In simulation on held-out objects from the filtered CMapDataset, we achieves average success rates of 92.40% on three seen hands and 74.02% on four unseen hands. In real-world experiments, the same model achieves an overall success rate of 76.00% on two unseen hands without additional training or fine-tuning. These results demonstrate the effectiveness of FunCo-Grasp in transferring grasp knowledge to unseen hands.
Figures & tables
Fig. 1: Overview of FunCo-Grasp . Top: The canonical functional schema and cross-embodiment functional correspondences established through a one-time lightweight functional annotation via (a) Functional Part Alignment and (b) Canonical Frame Alignment. Bottom: Grasp transfer from seen to unseen hands.
Fig. 2: Overview of FunCo-Grasp . (1) Cross-Embodiment Functional Correspondence maps physical links to shared functional parts according to their grasping roles and expresses them in canonical local frames. (2) Hand and Object Encoding extracts node features N , edge features E , and object patch features O . (3) Functional-Part-Based Grasp Generation conditions on the aligned hand representation and object geometry to generate the target spatial arrangement of the functional parts in the shared interaction space, before converting them into executable joint configurations via inverse kinematics.
Fig. 3: MAF Graph Encoder. Given the URDF and hand graph, the node branch encodes part geometry, grasping roles, and kinematics into N . The edge branch encodes reference part poses and kinematic connections into E .
TABLE I: Grasp success on held-out objects and per-grasp runtime. Red and yellow mark the best and second-best values. † UniMorphGrasp results come from its paper because its code is not publicly available. – denotes unavailable results.
Fig. 4: Grasps generated by FunCo-Grasp on held-out objects with seen (left) and unseen (right) hands. The same model transfers to unseen hands without target-hand grasp data, fine-tuning, or learned retargeting.
Fig. 5: Effect of FPA and CFA on ApexHand grasps: without FPA, hand regions are misplaced; without CFA, coordinated closure is disrupted.
Representation
FPA
CFA
Success Rate (%) ↑
Seen Hands
Unseen Hands
Shadow
Allegro
Barrett
Avg.
ApexHand
XHand
LeapHand
Robotiq-3F
Avg.
Full FunCo-Grasp
✓
✓
93.80
97.20
86.20
92.40
68.20
72.73
80.40
74.73
74.02
w/o FPA
✗
✓
94.20
93.00
85.00
90.73
13.40
45.20
9.20
67.80
33.90
w/o CFA
✓
✗
91.60
92.00
86.80
90.13
10.40
17.00
20.20
30.00
19.40
w/o FPA & CFA
✗
✗
89.80
94.60
87.80
90.73
3.00
0.00
0.80
0.00
0.95
TABLE II: Effect of Functional Part Alignment (FPA) and Canonical Frame Alignment (CFA) on grasp success across seen and unseen hands. Best results in each column are bold.
Fig. 6: Grasp success under finger-length scaling for seen (left) and unseen (right) hands, without retraining. 1.0× denotes original geometry.
TABLE III: Diversity of successful grasps on seen hands (rad.). Baselines use paper-reported values under their published protocols; † UniMorphGrasp uses its full model.
Fig. 7: Real-world setups with Revo2 (left) and ApexHand (right) mounted on UR5 arms, using RealSense D435i observations. Both hands are unseen during training.
Hand
Ball
Block
Bread
Cracker
Drill
ApexHand
10/10
10/10
9/10
9/10
7/10
Revo2
9/10
8/10
9/10
6/10
5/10
Vase
Milk
Mustard
Pepper
Soup
ApexHand
5/10
8/10
8/10
10/10
8/10
Revo2
4/10
8/10
6/10
8/10
5/10
TABLE IV: Per-object real-world grasp success on ApexHand and Revo2, both unseen during training. Each entry reports successful grasps out of ten trials.
Fig. 8: Representative successful grasps on Revo2 (top) and ApexHand (bottom). The same model is deployed on both hands without additional training or fine-tuning.
We study cross-embodiment 6-DOF robot grasping. Unlike prior works, we require the model not only to generalize to novel objects / scenes but also to novel gripper morphologies and physical grasping processes. Our method extends diffusion model based generative 6-DOF grasping models to condition on the additional gripper's representation. We propose a swept-volume heuristic for encoding the gripper. We train our cross-embodiment model with procedural grippers and a large-scale dataset of 2 Billion grasps. In simulation experiments, our model has the best zero-shot generalization to novel real-world grippers and objects over baseline methods. Our model also serves as a good initialization for fine-tuning to adapt to novel grippers. In ablations, we demonstrate the efficiency of our sweep-volume gripper representation and our procedural gripper training dataset. Last, we show zero-shot generalization to real-world novel grippers for 6-DOF grasping, surpassing baselines in cross-embodiment generalization.
Dexterous grasp generation across robot hands is challenging because hands differ in kinematic topology, actuation dimensions, and native command spaces. We introduce GraspGraphNet, a topology-aware grasp generation framework that represents each hand as a URDF-derived kinematic graph and directly generates executable palm poses and joint configurations. GraspGraphNet combines hierarchical object surface encoding, differentiable forward kinematics, and dynamic world-edge message passing to model evolving robot-object interactions. It applies conditional flow matching directly in executable palm-pose and joint-state space, avoiding post-processing optimization, inverse kinematics, and retargeting. Using a shared model trained on Barrett Hand, Allegro Hand, and Shadow Hand, GraspGraphNet achieves an average success rate of 83.48% with 40ms inference time per grasp on a 40-object benchmark. Without retraining, the same model achieves 72.70% success on controlled finger-removal variants, demonstrating robustness to hand-topology variations. These results suggest that graph-structured hand representations can effectively support dexterous grasp generation across robot hands with different kinematic structures. Project: https://lysees.github.io/graspgraphnet-page
Yeonseo Lee, Taeyeop Lee, Hyosup Shin +2
Korea Advanced Institute of Science and Technology (KAIST), Daejeon, South Korea.
Cross-embodiment dexterous grasping aims to synthesize stable grasps across heterogeneous multi-fingered hands with little or no embodiment-specific tuning. Existing interaction-centric methods achieve promising results, but their object representations often underrepresent local surface geometry, while their robot descriptors do not explicitly encode both robot morphology and kinematics. We propose MANGO-Grasp, an anisotropic interaction framework that represents objects as geometry-oriented 3D Gaussian primitives and robot hands as surface keypoints encoded into morpho-kinematic descriptors. The object primitives are adaptively allocated by geometric complexity and shaped as surface-aligned plates with outward normals, encoding local geometry. Mahalanobis fields over keypoint--primitive pairs serve as interaction prediction targets during training and as optimization guidance for grasp realization at inference. These fields rise sharply for displacement along the surface normal but only gently within the tangent plane, matching the directional structure of contact. Grasps are realized with one shared optimization formulation and hyperparameter setting across all embodiments. On the CMAP and MultiGripperGrasp benchmarks, MANGO-Grasp outperforms the strongest seen-hand baseline by up to 8.24 percentage points in simulation. It also transfers zero-shot to the unseen SharpaWave hand, improving over the strongest zero-shot baseline by up to 16.57 percentage points, and achieves 86% success in real-world experiments. The code and additional materials will be made available upon publication at https://connor-zh.github.io/MANGO-Grasp/.
Heng Zhang, Kevin Yuchen Ma, Mike Zheng Shou +2
Robotics & Autonomous Systems Division, Institute for Infocomm Research, Agency for Science, Technology and Research (A*STAR-I2R), Singapore · College of Computing and Data Science, Nanyang Technological University, Singapore · Show Lab, National University of Singapore, Singapore