Organizations: ACE Robotics · CUHK, MMLab · CPII under InnoHK · Tongji University · Shanghai Jiao Tong University · Zhejiang University · Nanyang Technological University
Dexterous grasping is the foundational primitive in embodied AI, demanding massive data to train robust models. As real-world data collection is expensive, simulation has become the mainstream paradigm. Yet, while cluttered scenes best reflect real-world applications, learning to grasp within them is bottlenecked by a critical scarcity of large-scale data. To resolve this, we curate high-quality 3D objects and supporting bases, proposing a scalable seed-and-filter strategy that bypasses sluggish scene-level optimization. This yields an unprecedented benchmark comprising over 2.6 million scenes and 0.4B scene-specific grasp ground truths, featuring diverse realistic layouts paired with rich semantic and geometric observations. Furthermore, we introduce the OmniDex model to overcome the grasp multimodality and last-millimeter precision errors plaguing current generative models. By coupling Soft Winner-Takes-All learning with human-inspired physical constraints during training, and utilizing physics-driven ranking, our approach achieves robust dexterous grasping without the latency of post-optimization. Experimental results show that OmniDex model achieves state-of-the-art performance and strong generalization across diverse scenes, views, and unseen objects.
Figures & tables
Figure 1 : Overview of OmniDex , the largest benchmark for dexterous grasping in cluttered scenes to date. Left : Multi-view RGB-D and instruction masks paired with scene-specific grasp poses. Right : A scalable seed-and-filter strategy enables unprecedented scale (2.6M scenes, 0.4B grasp ground truths), encompassing diverse objects, supporting bases, and realistic layouts.
Method
# Scenes
Setting
# Objects
Workspace
Modality
Dex1B [ 31 ]
-
Single
6K
-
Depth
AffordDexGrasp [ 27 ]
2K
Single
1K
Table
RGB-D
DDGC [ 16 ]
0.4K
Cluttered
0.3K
Table
RGB-D
DexGraspNet 2.0 [ 32 ]
8K
Cluttered
1K
Table
Depth
ClutterDexGrasp [ 5 ]
1K
Cluttered
2K
Table
Depth
Ours
2.6M
Cluttered
6K
Table, Box, Shelf
RGB-D
Table 1 : Quantitative comparison of simulation-based dexterous grasping benchmarks. Our dataset features the large scale of scenes and objects, emphasizing diverse layouts in clutter.
Figure 2 : Seed-and-filter generation pipeline . Our two-stage framework generates scene-valid grasps by: (1) constructing a reusable seed grasp library in free space, and (2) simulating cluttered scenes in Isaac Gym for collision filtering, followed by multi-modal rendering in Blender.
Figure 3 : Dataset statistics and visualization of 3D assets . (a): Instance counts and category-wise distributions of the curated graspable objects, highlighting the top 50 semantic classes. (b): Visual examples of the assets, which range from diverse grasped objects (medicine bottles, boxes, apples) to realistic supporting bases (tables, open-top boxes, shelves).
Subset
Obj. Div.
Scene Obj. Comp.
Sup. Base Type
Sup. Base Div.
Density
Scene Num.
Generalization Split
Train
Seen
Unseen
2 Dstd
✓
Multi-category
Table
✗
0.1 / 0.2 / 0.3
1.14M
919K
62K
163K
Dmix
✓
Multi-category
Table
✓
0.1 / 0.2 / 0.3
0.94M
706K
80K
152K
Dbox
✓
Multi-category
Box
✓
0.9
0.14M
112K
12K
16K
Dgrid
✓
Single-category
Table
✗
-
0.28M
246K
11K
24K
Dshelf
✓
Multi-category
Shelf
✓
-
0.14M
107K
11K
28K
Table 2 : Benchmark diversity and evaluation protocol across OmniDex subsets. Each subset is organized with a large training split and two decoupled test tracks to evaluate generalization.
Figure 4 : Architecture of the OmniDex Model . Condition embeddings are transformed from RGB-D inputs, instruction masks, and scene parameters. During training, a denoising network adds identical noise to M grasp poses, optimized jointly via a Soft Winner-Takes-All loss (for multimodal distributions) and a physics-aware loss (for stable contact). During inference, the network generates M candidates, followed by a physics-based ranking module to select the optimal grasp.
Method
Subset
Input View
Seen Objects
Unseen Objects
SR scene
SR scene∗
SR any
SR any∗
SR one
SR one∗
CFR
SR scene
SR scene∗
SR any
SR any∗
SR one
SR one∗
CFR
Ours
Dstd
Ego
33.60
55.80
18.87
36.79
65.15
79.35
84.12
40.15
59.74
27.18
40.77
65.54
81.77
84.61
Dmix
Ego
32.05
52.21
17.12
34.04
58.51
74.52
73.12
37.90
55.10
24.13
35.19
62.41
78.02
83.22
Dbox
Ego
32.34
35.83
17.06
19.79
55.15
56.52
99.12
33.48
36.50
17.80
20.17
56.77
58.11
95.72
Dgrid
Ego
55.01
79.60
33.68
55.02
91.91
93.70
95.46
43.11
68.05
33.97
52.67
87.50
92.09
94.86
Dshelf
Ego
29.49
39.46
18.31
25.25
55.17
60.57
81.95
14.96
24.08
11.58
15.08
34.26
37.89
63.57
Table 3 : Main Results on the OmniDex Benchmark . We evaluate our method against baselines across diverse scene subsets and compare with other methods on Dstd . * denotes the use of the finger-force trick [ 24 ]
Metrics
20%
40%
60%
80%
100%
SRscene
40.50
41.04
43.84
44.32
44.41
CFR
81.54
82.64
82.13
84.15
83.66
Table 4 : Quantitative results when trained on incrementally larger fractions of training dataset.
RGB
Depth
Cam. Para.
Table Height
Multi-view
SRscene
CFR
✓
✓
✓
31.89
79.97
✓
✓
✓
36.80
82.81
✓
✓
✓
43.68
84.09
✓
✓
✓
40.64
84.04
✓
✓
✓
✓
✓
49.87
82.11
✓
✓
✓
✓
44.41
83.66
Table 5 : Quantitative ablation on various input conditions. The highlighted row represents baseline.
Training Phase
Inference Phase
Metrics
Lstand
LWTA
LSDF
Lthumb,Lopp
Lclosure
Naive Samp.
Rank.
TTP
SRscene
CFR
✓
✓
37.79
81.37
✓
✓
38.07
81.31
✓
✓
✓
41.93
84.06
✓
✓
✓
✓
43.19
84.40
✓
✓
✓
✓
✓
36.08
79.71
Table 6 : Component-wise ablation study evaluated on both training and inference phases. TTP denotes Test-Time Optimization.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Metric
Parallel Hypotheses ( M )
TTP
1
2
4
8
16
32
64
SRscene
36.08
38.02
41.06
42.47
43.41
44.37
44.41
36.16
CFR
79.71
80.81
81.58
82.26
82.62
83.09
83.66
2.56
Inference Time (ms)
219.14
233.35
266.39
325.20
543.93
948.33
1598.41
6044.60
Appendix
Table D1 : Effect of parallel hypothesis generation vs. Test-Time Optimization.
SRscene failures
CFR failures
Failure type
Rate
Failure type
Rate
Insufficient contact
4.29%
Supporting-base collision
61.26%
Stability failure
95.71%
Non-target-object collision
38.74%
Appendix
Table D2 : Failure breakdown on the Dstd mini-split.
Rear offset (cm)
Height offset (cm)
IK feasibility (%)
0.0
+2.5
87.98
0.0
+5.0
87.94
2.5
+5.0
86.68
5.0
+2.5
85.63
7.5
+5.0
83.75
Appendix
Table D3 : IK feasibility under five representative Franka arm placements on the Dstd mini-split.
Figure E1 : Qualitative visualization of the Dstd subset.
Figure E2 : Qualitative visualization of the Dmix subset.
Figure E3 : Qualitative visualization of the Dbox subset.
Figure E4 : Qualitative visualization of the Dgrid subset.
Figure E5 : Qualitative visualization of the Dshelf subset.
Figure E6 : Qualitative results of OmniDex model across diverse benchmark subsets.
Figure E7 : Representative real-robot grasping trials in cluttered scenes.
This work addresses sequentially grasping multiple objects with a single dexterous hand without releasing those already held. Most dexterous grasping methods commit all of the hand's degrees of freedom to a single object, underutilizing its dexterity and leaving no redundancy for subsequent grasps. The proposed solution, MoDex, is a diffusion policy that predicts the next gripper pose directly from observations, conditioned on an opposition space and point cloud. The opposition space condition specifies which fingers participate in the current grasp, enabling the gripper to use only a subset of its available degrees of freedom while reserving the remaining degrees of freedom for subsequent grasps. To facilitate sim-to-real transfer, MoDex is trained in two stages: first through imitation learning on expert demonstrations, and subsequently through reinforcement learning fine-tuning, which consistently improves success rates over the pre-trained policy. We evaluate MoDex in simulation on a MuJoCo-based Franka Emika Panda robot equipped with an Allegro Hand and on the corresponding real-world hardware platform. Across both simulation and real-world experiments, MoDex achieves higher success rates than the evaluated learning-based baselines, improving performance by 2.92-17.92% and 6.67-17.78%, respectively. Project page: https://modex2026.github.io/.
Haofei Lu, Hongjia Liu, Yifei Dong +3
Department of Robotics, Perception and Learning, KTH Royal Institute of Technology, Sweden. · Robotics and Autonomous Systems at University of Turku, Finland.
Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requirements. Vision-language models (VLMs) offer a natural way to specify these requirements using language, but existing approaches either use a VLM to predict the grasp directly with limited spatial awareness, or train the VLM together with the grasping model, which requires significantly more data and compute. These limitations impede performance and have prevented scaling to multiple embodiments in complex scenes. We address this by proposing SeededGrasp, a novel data-efficient framework that enables a VLM to predict a seed point to be used as conditioning for a subsequent lightweight grasp-generation model. Our architecture decouples high-level semantic reasoning from low-level geometric execution, enabling multi-embodiment support while bypassing the need for expensive end-to-end training. To enable training such models, we release the first multi-embodiment tabletop grasping dataset comprising over 2.5M grasps in cluttered scenes. Experimental results demonstrate that our approach outperforms existing baselines, achieving 72% success in simulation and 78% in real-world grasping experiments. See our project site for data and code: https://uoft-isl.github.io/seeded-grasp/
Yang Xu, Gurpreet Singh Mukker, Raymond Wang +3
University of Toronto · Vector Institute · University of British Columbia +1
Learning robust dexterous grasping requires real-world data that records the physical outcomes of grasp attempts. Such data is hard to obtain at scale: teleoperation yields valid physical outcomes but is slow and operator-biased, while simulation-based generation is cheap and scalable but cannot certify contact validity. A natural solution is to generate candidate grasps and verify them on real hardware, but this scales only if the entire collection loop (perception, execution, labeling, and reset) runs without human intervention. We present AutoDex, an automated real-world data-collection system that closes this loop: for each candidate from a replaceable generator, it localizes the object under severe hand-object occlusion with dense 20-camera perception, executes collision-monitored robot motions, labels lift-and-hold success or failure, and actively resets the object between trials to expose additional candidates across stable poses. The result is a reusable database of physically labeled grasp trials that downstream systems can query by retrieval and feasibility filtering. Using AutoDex, we collect 3,593 grasp trials across Allegro and Inspire hands on 100 diverse objects, with synchronized multi-view observations and robot-state logs. For a matched 500-trajectory collection, AutoDex requires 10.3 h versus 49.4 h for teleoperation, yielding a 4.8x throughput improvement, and grasps retrieved from the AutoDex-validated database succeed 76% versus 34% for simulation-only validation. Code and data will be publicly released.