cs.ROSep 3, 2026

Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis

Authors: Sixu Yan, Shikang Wang, Binhua Huang, Xuanlai Tang, Guohua Fan, Fan Huang, Haoxuan Li, Yongkang Li, +6 more

Organizations: School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan 430074, China · Hubei Automation Institute, Wuhan 430071, China · KEENON Robotics Co., Ltd., Shanghai 201206, China · Suzhou Silicon Era Intelligent Technology Co., Ltd., Suzhou 215131, China · Suzhou Zhichuang Xinwei Technology Co., Ltd., Suzhou 215123, China · College of Materials, Xiamen University, Xiamen 361005, China · School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China · ByteDance, Beijing 100098, China · State Key Laboratory of General Artificial Intelligence, Beijing Institute for General Artificial Intelligence (BIGAI), Beijing 100080, China · School of Artificial Intelligence, Peking University, Beijing 100871, China

Abstract

This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizable grasp synthesis across different robotic hands. Unlike existing VLG methods that tightly couple foundation models with end-to-end grasp policies, AdaRoboVLG learns an efficient generalizable base policy that generates and evaluates physically feasible grasp candidates through explicit kinematic mapping and force-closure-based stability estimation, while offloading task-dependent understanding to specialized foundation-model modules. These modules provide composable priors that are integrated into the grasp synthesis process, enabling contextually adaptive grasp synthesis without retraining the underlying grasp policy. Through extensive simulation and real-world experiments, we demonstrate that (i) the base policy exhibits efficient learning and strong cross-hand generalization, (ii) the framework effectively incorporates spatial, cognitive, and temporal priors to address three representative grasping challenges without compromising grasp synthesis performance compared to state-of-the-art methods, and (iii) these priors can operate jointly to enable functional grasping in cluttered and dynamic environments. These results indicate that decoupling physical grasp synthesis from task-dependent understanding provides a scalable paradigm for robotic grasping, allowing future advances in foundation models to be directly translated into improved grasp capabilities without redesigning or retraining the underlying grasp policy. Supplementary videos are available at https://adarobovlg.github.io/

Explore similar work

CardsList