cs.ROSep 30, 2026

HiWE: Hierarchical World Knowledge Model with Visual Keypoint Enhancement for Zero-Shot 3D Path Planning

Authors: Guoqing Ma, Mingqi Yuan, Chen Gao, Jiayu Chen, Shan Yu

Organizations: Institute of Automation, Chinese Academy of Sciences, Beijing, China · School of Future Technology, University of Chinese Academy of Sciences · State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, CAS · The University of Hong Kong, Hong Kong SAR, China · Department of Electronic Engineering, Tsinghua University, China

Abstract

Robot demonstration generation requires a system to identify where an interaction should occur, plan a feasible motion, and execute the required contact. HiWE connects these decisions through a point-based interface between visual grounding and language-based planning. PointVLM is instruction-tuned to associate task-relevant objects with image coordinates using a mixture of point annotations, segmentation-derived samples, robot observations, and visual question answering data. Depth measurements lift these predictions into a semantic 3D representation. A language planner, 3DLLM, uses this representation to specify end-effector waypoints and gripper commands, while a hybrid grasping module resolves local grasp poses. The evaluation covers 14 simulated manipulation tasks and four physical-robot tasks, together with ablations of the visual training data, spatial inputs, and grasp selection. Here, zero-shot execution refers to deployment without task-specific demonstration training; the visual model uses existing robot data during fine-tuning. This paper describes the original point-based formulation of the framework; its relationship to the subsequent GeneralVLA extension is detailed in the introduction.

Figures & tables

Explore similar work

CardsList
  1. Grounded World Model: Latent Planning with Language Goals

    Apr 13, 2026Quanyi Li, Lan Feng, Haonan Zhang +4World ModelsWorld

  2. ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning

    Jun 17, 2026Jisoo Kim, Sangwon Baik, Taeksoo Kim +43D Visual GroundingRgb-Depth Cameras

  3. 3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints

    Sep 8, 2026Ziqin Huang, Yingyue Li, Chenyangguang Zhang +6WaypointsRobotic Manipulation