Most current grasp synthesis systems are trained offline and remain fixed during deployment. While this works well when deployment conditions resemble the training data, performance can degrade when robots encounter conditions they have not seen before, such as unfamiliar objects. In this work, we present a continual-learning framework for single-view 6-DoF grasp synthesis for a parallel-jaw gripper in cluttered scenes. Rather than finetuning a large parametric model, our method adapts through memory in a learned embedding space: grasp outcomes update future grasp scores, while optional user demonstrations are recalled and transferred to new scenes as additional candidate grasps. We evaluate our method in simulation and in extensive real-world experiments comprising over 1500 grasp trials. We show that our method matches the performance of existing 6-DoF grasping baselines even before adaptation, improves online on unseen objects from categories absent or underrepresented during training, and supports long-horizon continual learning with limited forgetting. In real-world experiments, our method reaches over 90% success rates on several challenging object categories after only 50 online grasp attempts. Videos and code at https://giuschio.github.io/cl_grasping/.
Figures & tables
Figure 1: We present a continual-learning framework for 6-DoF grasp synthesis that is pretrained in simulation and keeps adapting during deployment. By leveraging memory in a learned embedding space, our method adapts from grasp outcomes and recalled user demonstrations without updating network weights.
Method
Train Obj.
Airplane
Animal
Bottle
Bowl
Drill
Fasteners
Hammer
Mug
Pliers
Screwdriver
Average (excl. training)
ICGNet
91.6%
75.4%
89.0%
95.4%
52.1%
95.8%
81.8%
88.1%
85.2%
56.3%
67.0%
78.6%
ContactGraspNet ———
89.2%
83.0%
89.5%
96.6%
95.4%
93.1%
90.6%
74.1%
81.6%
39.7%
74.5%
81.8%
EdgeGraspNet
95.3%
91.6%
89.0%
96.3%
91.0%
96.8%
91.8%
96.0%
91.2%
93.1%
92.6%
92.9%
Ours (base)
98.1%
92.8%
93.8%
96.4%
88.6%
93.5%
97.4%
96.0%
92.4%
98.0%
97.2%
94.6%
Table 1: Grasp success rates without online adaptation (simulation). For each object category and method, results are computed over 2,500 grasp attempts. Our base model performs better than or comparably to the baselines across all categories, while achieving the highest average performance.
Method
Train Obj.
Airplane
Animal
Bottle
Bowl
Drill
Fasteners
Hammer
Mug
Pliers
Screwdriver
Average (excl. training)
Ours (base)
98.1%
92.8%
93.8%
96.4%
88.6%
93.5%
97.4%
96.0%
92.4%
98.0%
97.2%
94.6%
Ours (+ Airplane)
97.8%
97.0%
-
-
-
-
-
-
-
-
-
-
Ours (+ Animal)
98.1%
96.2%
95.7%
-
-
-
-
-
-
-
-
-
Ours (+ Bottle)
98.9%
96.2%
95.0%
99.5%
-
-
-
-
-
-
-
-
Ours (+ Bowl)
98.7%
96.8%
96.0%
99.5%
98.9%
-
-
-
-
-
-
-
Ours (+ Drill)
98.7%
96.0%
95.0%
97.5%
98.3%
97.9%
-
-
-
-
-
-
Table 3: Sequential continual learning in simulation. The robot adapts to the ten evaluation categories one after another. After each stage, we evaluate on the current category and all previously seen categories, using 1000 grasp attempts per category. Our method adapts across categories without forgetting previous ones. Final performance on average remains close to category-specific adaptations (97.8% vs. 98.1%, see Table 2 ).
Figure 2: Real-world setup and object sets used for evaluation. The control objects are chosen to resemble common grasp-synthesis evaluation objects, while the remaining categories include mugs and bowls, kitchen tools, pliers, screwdrivers, and toys with more challenging geometries, materials, thin structures, and non-uniform mass distributions. The single-camera scenes have moderate occlusion. Stacked or interlocked clutter is outside our scope.
Figure 3: Qualitative examples of real-world adaptation. On the pliers, repeated outcomes shift grasp preferences away from the low-friction handles and toward a grasp that locks against the metal head. On the ladle, they shift grasps closer to the center of mass. On the shampoo bottle, the pointcloud misses much of the object body from the thin side, so the geometric sampler does not reliably generate grasps that wrap around the wider side. Demonstration recall adds this missing grasp mode to the proposal set.
Single-category adaptation
Continual learning
Method
Control objects
Mugs and Bowls
Kitchen Tools
Pliers
Screwdriver
Toys
Method
Mixed scenes
M2T2
53.8% (65.0%)
82.5% (100%)
32.4% (35.0%)
11.0% (5.0%)
11.3% (10.0%)
77.4% (90.0%)
M2T2
50.0% (45.0%)
AnyGrasp
78.8% (95.0%)
83.3% (95.0%)
71.9% (90.0%)
28.6% (30.0%)
54.9% (90.0%)
85.0% (100%)
AnyGrasp
69.2% (80.0%)
EdgeGraspNet
80.2% (90.0%)
88.5% ( 100% )
37.8% (20%)
0% (0%)
0% (0%)
72.5% (85.0%)
EdgeGraspNet
60.0% (30.0%)
Ours (base)
80.0% (95.0%)
92.7% ( 100% )
77.4% (95.0%)
62.5% (80.0%)
85.7% (90.0%)
83.6% ( 100.0% )
Ours (base)
72.4% (85.0%)
Ours (w adaptation)
93.2% ( 100% )
100% ( 100% )
91.3% ( 100.0% )
68.3% ( 90.0% )
96.0% ( 100.0% )
94.4% ( 100.0% )
Ours (merged)
89.6% ( 100.0% )
Table 4: Real-world evaluation. Left block: success rate and scene-clear rate (in parentheses) on each real-object category, comparing M2T2, AnyGrasp, and EdgeGraspNet against our base model and our model after independent category-level adaptation. Right block: performance on scenes containing objects from all categories, using a single model adapted to all categories. Online adaptation improves our model on every category and exceeds 90% success rate on five of six categories. On mixed scenes, the continual-learning model reaches 89.6% success and 100.0% scene-clear rate. In total, our real-world experiments comprise more than 1500 grasp attempts.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Split/category
IID grasps
Airplane
Animal
Bottle
Bowl
Drill
Fasteners
Hammer
Mug
Pliers
Screwdriver
Median 5-NN distance
0.36
0.54
0.56
0.38
0.48
0.43
0.27
0.48
0.59
0.34
0.41
Appendix
Table 5: Median fifth-nearest-neighbor distance from each category’s grasp embeddings to the offline training memory. IID grasps consists of held-out grasps on training objects. Larger values indicate grasp-local geometry farther from the offline training memory.
Figure 4: Grasp-frame convention used throughout the paper. The grasp frame is placed at the center of the two fingers. The y axis is defined by the finger closing direction. The z axis is the approach direction.
Figure 5: Failure cases for contact-normal grasp proposals on thin objects. The left panel shows the actual scene geometry: a thin object with vertical sides resting on the table. In real depth observations, these small side surfaces may be smoothed into slopes or absent from the pointcloud. The contact-based sampler then either follows an invalid sloped normal, causing finger-table collision, or has no side contacts from which to sample the grasp.
Figure 6: Local grasp patch extraction and Basis Point Set (BPS) encoding. Given a candidate grasp, we crop the local point cloud around the gripper, express it in the grasp frame, and encode it with fixed BPS basis points. For clarity, the figure visualizes a 2D slice of the BPS encoding. In the full representation, each basis point stores the displacement to its nearest observed scene point together with that point’s surface normal.
Figure 7: Open3D interface used to provide demonstrations. The interface displays the observed point cloud and an editable gripper model. Users adjust the gripper pose with keyboard controls and confirm the demonstration by pressing enter.
Method
Train Obj.
Airplane
Animal
Bottle
Bowl
Drill
Fasteners
Hammer
Mug
Pliers
Screwdriver
Average (excl. training)
Ours (base) (BCE emb)
-
91.9%
94.3%
96.9%
93.8%
95.7%
96.8%
95.2%
90.6%
98.1%
93.8%
94.7%
Ours (+ Airplane)
-
95.9%
-
-
-
-
-
-
-
-
-
-
Ours (+ Animal)
-
94.7%
94.1%
-
-
-
-
-
-
-
-
-
Ours (+ Bottle)
-
94.2%
95.5%
99.0%
-
-
-
-
-
-
-
-
Ours (+ Bowl)
-
95.3%
95.7%
99.1%
98.6%
-
-
-
-
-
-
-
Ours (+ Drill)
-
93.2%
95.0%
99.3%
98.3%
97.9%
-
-
-
-
-
-
Appendix
Table 6: Sequential continual learning with a BCE embedding. The model uses the same memory-based update rule as the main method, but replaces the SNN-trained embedding with embeddings from a binary classifier.
Figure 8: PCA cumulative-variance comparison of the learned embedding spaces. Each plot shows the cumulative explained variance as a function of the number of principal components. The BCE embedding concentrates most variance in the first component, while the SNN+reconstruction embedding distributes variance across more dimensions, suggesting that it more effectively uses the embedding space to preserve geometry and separate multiple success and failure modes.
Method
Train Obj.
Airplane
Animal
Bottle
Bowl
Drill
Fasteners
Hammer
Mug
Pliers
Screwdriver
Average (excl. training)
Ours (base)
98.1%
92.8%
93.8%
96.4%
88.6%
93.5%
97.4%
96.0%
92.4%
98.0%
97.2%
94.6%
Ours (+ Airplane)
97.8%
97.0%
96.2%
98.8%
96.3%
93.4%
96.5%
99.2%
92.0%
99.3%
99.2%
96.8%
Ours (+ Animal)
98.1%
96.2%
95.7%
97.8%
97.8%
93.1%
97.1%
98.5%
91.6%
98.5%
99.0%
96.5%
Ours (+ Bottle)
98.9%
96.2%
95.0%
99.5%
94.0%
94.0%
99.2%
99.2%
94.3%
98.3%
98.8%
96.8%
Ours (+ Bowl)
98.7%
96.8%
96.0%
99.5%
98.9%
96.1%
98.4%
99.1%
95.4%
98.7%
98.6%
97.8%
Ours (+ Drill)
98.7%
96.0%
95.0%
97.6%
98.3%
97.9%
98.9%
98.9%
94.3%
98.9%
98.9%
97.5%
Appendix
Table 7: Forward adaptation in simulation. The robot adapts to the ten evaluation categories one after another. After each stage, we evaluate on the current category and all categories (seen and unseen), using 1000 grasp attempts per category.
Method
Mem. size (MB)
Sampling (ms)
Recall (ms)
Collision Check (ms)
Encoder (ms)
Scorer (ms)
Ours (base)
1.20 MB
630 ms
-
10 ms
65 ms
50 ms
Ours (end of adaptation)
1.35 MB
630 ms
700 ms
25 ms
80 ms
60 ms
Appendix
Table 8: Runtime and memory before and after the sequential adaptation experiment. Timings are measured on a machine with 32 GB RAM, an NVIDIA GeForce RTX 4090, and an Intel Core i9-12900K.
Object set
Recall
Geometric sampling
Control objects
66%
34%
Mugs and bowls
0%
100%
Kitchen tools
59%
41%
Pliers
10%
90%
Screwdrivers
68%
32%
Toys
39%
61%
Appendix
Table 9: Source distribution of executed grasps after real-world adaptation. Values report the percentage of executed grasps originating from demonstration recall and geometric sampling.
Figure 9: Objects used for the real-world demonstration comparison. The objects were selected from real-world cases that were difficult for the robot before adaptation.
Figure 10: Real-world adaptation traces with and without demonstrations. Each panel shows 20 adaptation attempts on one object, starting from the same base model. Bars indicate failed grasp attempts; missing bars indicate successful attempts. The black trace adapts only from binary grasp outcomes, while the blue trace uses the full method with demonstrations. For the full method, the demonstration memory is initialized with two demonstrations, and a new demonstration is added after each failed grasp.
This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizable grasp synthesis across different robotic hands. Unlike existing VLG methods that tightly couple foundation models with end-to-end grasp policies, AdaRoboVLG learns an efficient generalizable base policy that generates and evaluates physically feasible grasp candidates through explicit kinematic mapping and force-closure-based stability estimation, while offloading task-dependent understanding to specialized foundation-model modules. These modules provide composable priors that are integrated into the grasp synthesis process, enabling contextually adaptive grasp synthesis without retraining the underlying grasp policy. Through extensive simulation and real-world experiments, we demonstrate that (i) the base policy exhibits efficient learning and strong cross-hand generalization, (ii) the framework effectively incorporates spatial, cognitive, and temporal priors to address three representative grasping challenges without compromising grasp synthesis performance compared to state-of-the-art methods, and (iii) these priors can operate jointly to enable functional grasping in cluttered and dynamic environments. These results indicate that decoupling physical grasp synthesis from task-dependent understanding provides a scalable paradigm for robotic grasping, allowing future advances in foundation models to be directly translated into improved grasp capabilities without redesigning or retraining the underlying grasp policy. Supplementary videos are available at https://adarobovlg.github.io/
Sixu Yan, Shikang Wang, Binhua Huang +11
School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan 430074, China · Hubei Automation Institute, Wuhan 430071, China · KEENON Robotics Co., Ltd., Shanghai 201206, China +7
Dexterous grasp synthesis has advanced rapidly in generating stable and physically plausible hand poses, but real-world manipulation requires grasps that preserve the function implied by the task. We study open-vocabulary task-oriented dexterous grasp generation, where a robot must infer functional intent from free-form language, ground it in multi-view visual observations and object geometry, and generate an executable high-degree-of-freedom grasp. We present OpenDexGrasp, a unified data and generative modeling framework for this setting. OpenDexVerse provides dual-source supervision organized by the Coverage-to-Alignment (C2A) Recipe: OpenDex-Scale offers large-scale semantic and geometric coverage through automatic grasp synthesis and vision-language annotation, while OpenDex-Align supplies high-quality embodied alignment through human teleoperation and category-level transfer. OpenDexGrasp learns a shared perception-action latent representation that couples open-vocabulary vision-language context with dexterous action generation. Affordance grounding and grasp generation provide complementary supervision over this latent space, enabling direct generation of task-consistent dexterous grasps without a separate affordance-to-pose inference stage. Extensive simulation and real-robot experiments demonstrate improved functional alignment, physical feasibility, generalization to unseen categories, and real-world execution success. Additional details and videos are available at https://opendexgrasp.github.io/.
Jiyao Zhang, Junhan Wang, Tianyu Wang +4
CFCS, School of CS, PKU, China · National Key Laboratory for Multimedia Information Processing, School of CS, PKU, China · PrimeBot +1
Dexterous grasping across diverse object scales requires contact modes ranging from two-finger pinches to bimanual grasps. Existing dexterous grasp synthesis methods reduce the high-dimensional optimization space with manually designed expected contacts and initialization heuristics, which struggle to balance synthesis success rate and diversity. We present HUGS (Human-prior-guided Unified Dexterous Grasp Synthesis), a human-prior-guided framework for unified dexterous grasp synthesis across modes and scales. Instead of directly retargeting human demonstrations, HUGS learns an object-conditioned human prior that captures human grasp preferences and guides downstream force-closure-aware optimization. The prior is trained on a compact self-collected human grasp dataset with 1.8K grasps over 304 objects, providing broad coverage of object scales and contact modes. During synthesis, HUGS adaptively proposes contact modes and wrist initializations, substantially improving the balance between contact-mode coverage and synthesis success rate over heuristic-based methods. With HUGS, we synthesize 3.2M robotic grasps over 157K scenes, spanning object half-diagonal lengths from 2 cm to 30 cm and modes from two-finger to bimanual grasps. Models trained on the synthesized dataset autonomously select appropriate contact modes in the real world, enabling grasping from screws to large boxes.