CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning
Authors: Julien Merand, Boris Meden, Liming Chen, Mathieu Grossard
Organizations: Université Paris-Saclay, CEA, List, F-91120 Palaiseau, France · Ecole Centrale Lyon, CNRS, LIRIS, UMR5205, Institut Universitaire de France (IUF), F-69130 Ecully, France
Current dexterous grasp planners primarily optimize for physical stability, focusing on whether an object can be grasped rather than how it should be grasped to support downstream functional tasks. However, conditioning grasp synthesis on specific human grasp taxonomies typically requires prohibitively expensive, object-annotated datasets. To address these limitations, we propose CoToGrasp, a novel generative framework that synthesizes diverse, stable grasps strictly conditioned on specific contact topologies. To bypass the data collection bottleneck, CoToGrasp is trained entirely in an object-agnostic manner. We introduce a feature-based canonical workspace that projects local object features into a unified gripper-centric domain, effectively decoupling the semantic functional intent from the arbitrary object geometry. By learning the intrinsic contact manifold of the gripper within this workspace, our model achieves zero-shot generalization to unseen objects at inference. Extensive evaluations on the large-scale DexGraspNet dataset demonstrate that CoToGrasp achieves state-of-the-art performance, outperforming existing taxonomy-guided planners. Finally, we demonstrate the physical viability and kinematic feasibility of our synthesized contact topologies on a physical robot platform. Code is available on our project website at https://cea-list.github.io/cotograspweb/ .
Figures & tables
Figure 1 : Contact-Topology-Conditioned Grasp Synthesis. Given a desired semantic contact-topology condition (top left) – categorized into Precision , Object-Specific (highly constrained topologies tailored for specific tool use) or Power functional groups – and a novel, unseen object (bottom left), our framework synthesizes functionally diverse and physically stable grasps (right). Rather than learning contact topologies directly on the object geometry, we project local object features into a feature-based canonical workspace. This unified spatial representation effectively decouples the functional intent from the specific object identity. Within this workspace, we learn a latent manifold (center) that models the intrinsic contact capabilities of the gripper, enabling zero-shot generalization to diverse target geometries.
Figure 2 : CoToGrasp Method Overview. The proposed framework operates in two distinct phases. Top (Object-Agnostic Training): The model learns an intrinsic, gripper-centric contact manifold within a canonical feature-based workspace, independent of object geometry. Bottom (Grasp Synthesis): At inference, a target object is transformed into the canonical frame. The network’s contact-topology-conditioned prediction is strictly filtered through a validation pipeline before energy-based optimization aligns the gripper to yield the final stable grasp ( Q∗ ).
Figure 3 : Semantic Grasp Taxonomy and Contact Mapping. (Left) Correspondences between the classical Feix [ 8 ] (F) taxonomy (top), the haptic Gonzalez [ 9 ] (M) taxonomy (middle row) and our derived point cloud contact templates Am (bottom row). We categorized the 21 templates into three distinct functional groups: Precision , Power and Object-Specific (highly constrained topologies tailored for specific tool use). (Right) The 22 anatomical contact zones defined by Gonzalez (top) and the direct surjective mapping ( ζ ) onto our discrete gripper handprint H (bottom).
Method
Object-Agnostic Training
SR ↑
HTC↑
Speed (sec. / grasps)
Diversity (avg.) ↑
t (m)
R (rad)
Q (rad)
DFC [ 23 ]
✓
72.15
0.7389
>1800
0.0607
1.424
0.3579
GenDexGrasp [ 21 ]
✗
71.15
0.5956
14.65
0.0519
1.416
0.2567
DRO-Grasp [ 38 ]
✗
63.30
0.6504
1.72
0.0546
1.515
0.2892
GOAG [ 28 ]
✓
77.90
0.6527
0.20
0.0479
1.401
0.3170
CoToGrasp
✓
36.94
0.83
0.11
0.0674
1.4927
0.3458
Table 1 : Comparison with taxonomy-unaware baselines. CoToGrasp achieves the highest semantic entropy ( HTC ) and generation speed, overcoming the functional mode collapse typical of unconditioned planners.
Figure 4 : Functional contact topology distribution across taxonomy-unaware planners. The histogram illustrates the distribution of grasps generated by unconditioned baselines compared to CoToGrasp on the Multidex objects set. The unknown category represents physically stable grasps with unnatural contact patterns that fail to match any contact topology. Notably, unconditioned baselines exhibit a severe generative bias (mode collapse) toward enveloping power grasps (red box).
Method
SR (%)
HSR
TC (%)
HTC
Power
Precision
Obj. Spe.
Avg. Topo.
Avg. Obj.
Dexonomy [ 4 ]
27.16
12.36
19.62
21.13
23.80
0.91
14.28
0.77
CoToGrasp
29.75
22.71
25.50
26.72
27.56
0.96
17.18
0.84
CoToGrasp (w/o Label-Consistency)
25.11
14.77
20.85
21.14
22.97
0.94
14.45
0.81
CoToGrasp (w/o Force-Closure)
26.90
14.87
21.08
22.08
23.65
0.95
16.26
0.81
CoToGrasp (No Check)
25.09
14.73
20.58
21.06
23.00
0.94
14.72
0.81
Table 2 : Taxonomy-aware grasp synthesis. CoToGrasp outperforms the baseline in both physical stability (SR) and semantic accuracy (TC), particularly on highly constrained precision grasps.
Method
SR (%)
M2
M6
M11
M13
M18
M21
Dexonomy [ 4 ]
10.5
15.2
60.5 †
20.3
29.6
37.2 †
CoToGrasp
30.3
21.7
29.8
29.6
31.3
33.5
Table 3 : Per-Topology results: 3 Power grasps (M13, M18, M21), 2 Precision (M2, M6) and 1 Object-Specific (M11). † Indicates artificially inflated scores due to mode collapse, where Dexonomy defaults to unverified enveloping grasps.
Figure 5 : Real-World kinematic validation. CoToGrasp synthesizes diverse, topology-compliant grasps that are physically executable on a physical Allegro Hand using YCB [ 3 ] objects. The target contact topologies (indicated above each frame) demonstrate the physical viability of the generated grasps across both precision and power categories.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Handprint areas and workspace constitution. Left: Discretized handprints of the Shadow and Allegro hands, illustrating the manually defined anatomical zone divisions. Right: A three-quarter view of the Shadow Hand’s workspace.
Figure 7 : Taxonomy transfer mapping for anthropomorphic grippers. The handprints of the Shadow Hand (left) and the Allegro Hand (right) are segmented into manually defined, corresponding anatomical zones ( A1 – A21 ).
Module / Parameter
Notation
Value / Size
Training & Optimization
Batch Size
–
32
Learning Rate
–
10−5
Training Epochs
–
50
KLD Regularization Weight
β
0.1
Attention Factor
α
3.0
Appendix
Table 4 : Implementation Details and Network Hyperparameters for the CoToGrasp framework.
Raw DGCNN
Workspace
+ Attn.
Matched Pairs
0.35
0.61
0.66
Random Pairs
0.23
0.33
0.27
Appendix
Table 5 : Feature Alignment (Cosine Similarity).
Figure 8 : t-SNE visualization of features after canonical workspace projection.
Figure 9 : Topology-Conditioned Pose Sampling. For specific grasp topologies (such as M4 and M12), the initial object pose is sampled within a restricted kinematic region. Middle: The template’s full active sub-workspace is shown in light blue, while the truncated sub-workspace – filtered for reachability and palm clearance – is highlighted in red. Right: Examples of the initialized object point cloud O~ (green) successfully placed within this feasible region after the sampled spatial transformation.
Label
Eval.
HSR
TC (%)
HTC
Consistency
Isaac
✓
✓
0.96
17.18
0.84
✓
✗
21.92
0.53
✗
✓
0.94
14.45
0.81
✗
✗
19.34
0.50
Appendix
Table 6 : Topology compliance ablation.
Figure 10 : Illustration of metric-induced misclassification. While CoToGrasp strictly respects the target contact topology, the kinematic optimization may result in incidental contacts where an adjacent phalanx rests against the object surface. Despite the repulsive term in our optimization energy function, these phalanges often cannot be pushed away due to inherent kinematic constraints. For instance, the left grasp illustrates an intended M3 pinch reclassified as M12, while the right shows an M4 grasp reclassified as M15. Although the intended contacts are successfully achieved and the grasps remain physically stable, these incidental contacts trigger a strict reclassification by our automated pipeline.
Figure 11 : Histogram comparing the frequencies of effective topologies and attempted topologies (as defined Sec. 4.2 ) among all stable grasps generated by CoToGrasp (top) and Dexonomy [ 4 ] (bottom).
Object Complexity
Metric
CoToGrasp
Dexonomy [ 4 ]
Convex → Non-Convex
SR Retained
58.72%
37.42%
Severe Concavities ( c<0.4 , ∼4% data)
SR
18.30%
12.05%
TC
19.17%
8.94%
Appendix
Table 7: Object-Level Analysis. Evaluating performance across geometric complexity ( c=Vobj/Vhull ). CoToGrasp shows superior retention of both physical stability (SR) and semantic compliance (TC) on challenging non-convex objects.
Figure 12 : Qualitative Synthesis Gallery on the Allegro Hand. Synthesized grasp configurations for a diverse subset of the YCB object dataset.
Figure 13 : Topological Clustering of Shadow Hand Grasps. Grasps grouped by contact topology (M1-M21), demonstrating consistent semantic alignment across varied object classes.
Dexterous grasp synthesis has advanced rapidly in generating stable and physically plausible hand poses, but real-world manipulation requires grasps that preserve the function implied by the task. We study open-vocabulary task-oriented dexterous grasp generation, where a robot must infer functional intent from free-form language, ground it in multi-view visual observations and object geometry, and generate an executable high-degree-of-freedom grasp. We present OpenDexGrasp, a unified data and generative modeling framework for this setting. OpenDexVerse provides dual-source supervision organized by the Coverage-to-Alignment (C2A) Recipe: OpenDex-Scale offers large-scale semantic and geometric coverage through automatic grasp synthesis and vision-language annotation, while OpenDex-Align supplies high-quality embodied alignment through human teleoperation and category-level transfer. OpenDexGrasp learns a shared perception-action latent representation that couples open-vocabulary vision-language context with dexterous action generation. Affordance grounding and grasp generation provide complementary supervision over this latent space, enabling direct generation of task-consistent dexterous grasps without a separate affordance-to-pose inference stage. Extensive simulation and real-robot experiments demonstrate improved functional alignment, physical feasibility, generalization to unseen categories, and real-world execution success. Additional details and videos are available at https://opendexgrasp.github.io/.
Jiyao Zhang, Junhan Wang, Tianyu Wang +4
CFCS, School of CS, PKU, China · National Key Laboratory for Multimedia Information Processing, School of CS, PKU, China · PrimeBot +1
Dexterous manipulation requires planning a grasp configuration suited to the object and task, which is then executed through coordinated multi-finger control. However, specifying grasp plans with dense pose or contact targets for every object and task is impractical. Meanwhile, end-to-end reinforcement learning from task rewards alone lacks controllability, making it difficult for users to intervene when failures occur. To this end, we present GRIT, a two-stage framework that learns dexterous control from sparse taxonomy guidance. GRIT first predicts a taxonomy-based grasp specification from the scene and task context. Conditioned on this sparse command, a policy generates continuous finger motions that accomplish the task while preserving the intended grasp structure. Our result shows that certain grasp taxonomies are more effective for specific object geometries. By leveraging this relationship, GRIT improves generalization to novel objects over baselines and achieves an overall success rate of 87.9%. Moreover, real-world experiments demonstrate controllability, enabling grasp strategies to be adjusted through high-level taxonomy selection based on object geometry and task intent.
Juhan Park, Taerim Yoon, Seungmin Kim +10
Department of Artificial Intelligence, Korea University, Seoul, Korea · Korea University of Technology and Education, Cheonan, Korea · Naver AI Lab, Korea +4
Dexterous grasping across diverse object scales requires contact modes ranging from two-finger pinches to bimanual grasps. Existing dexterous grasp synthesis methods reduce the high-dimensional optimization space with manually designed expected contacts and initialization heuristics, which struggle to balance synthesis success rate and diversity. We present HUGS (Human-prior-guided Unified Dexterous Grasp Synthesis), a human-prior-guided framework for unified dexterous grasp synthesis across modes and scales. Instead of directly retargeting human demonstrations, HUGS learns an object-conditioned human prior that captures human grasp preferences and guides downstream force-closure-aware optimization. The prior is trained on a compact self-collected human grasp dataset with 1.8K grasps over 304 objects, providing broad coverage of object scales and contact modes. During synthesis, HUGS adaptively proposes contact modes and wrist initializations, substantially improving the balance between contact-mode coverage and synthesis success rate over heuristic-based methods. With HUGS, we synthesize 3.2M robotic grasps over 157K scenes, spanning object half-diagonal lengths from 2 cm to 30 cm and modes from two-finger to bimanual grasps. Models trained on the synthesized dataset autonomously select appropriate contact modes in the real world, enabling grasping from screws to large boxes.