Organizations: Institute of Automation, Chinese Academy of Sciences, Beijing, China · School of Future Technology, University of Chinese Academy of Sciences · State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, CAS · The University of Hong Kong, Hong Kong SAR, China · Department of Electronic Engineering, Tsinghua University, China
Robot demonstration generation requires a system to identify where an interaction should occur, plan a feasible motion, and execute the required contact. HiWE connects these decisions through a point-based interface between visual grounding and language-based planning. PointVLM is instruction-tuned to associate task-relevant objects with image coordinates using a mixture of point annotations, segmentation-derived samples, robot observations, and visual question answering data. Depth measurements lift these predictions into a semantic 3D representation. A language planner, 3DLLM, uses this representation to specify end-effector waypoints and gripper commands, while a hybrid grasping module resolves local grasp poses. The evaluation covers 14 simulated manipulation tasks and four physical-robot tasks, together with ablations of the visual training data, spatial inputs, and grasp selection. Here, zero-shot execution refers to deployment without task-specific demonstration training; the visual model uses existing robot data during fine-tuning. This paper describes the original point-based formulation of the framework; its relationship to the subsequent GeneralVLA extension is detailed in the introduction.
Figures & tables
Figure 1: Interfaces for robot manipulation. HiWE exposes task-relevant points and planned trajectories between perception and execution. This separation allows the visual component to be trained on annotations that do not require deployment-specific action trajectories.
Component
HiWE
GeneralVLA
Perception
PointVLM coordinate prediction
ASM segmentation and refinement
Planning
Current task and 3D scene points
3D scene planning with KnowledgeBank
Experience
No persistent retrieval in 3DLLM
Retrieval, construction, consolidation
Execution
HGM grasp selection
HGM grasp selection
Table 1: Scope of the original system and its subsequent extension. Shared components are not independent evidence of a new architecture.
Figure 2: HiWE data interfaces. PointVLM associates task-relevant objects with image points. Depth supplies their 3D coordinates for 3DLLM, which outputs waypoints and gripper commands. HGM provides local grasp poses for the execution stage. The overall decomposition and grasping component are also described in GeneralVLA ( Ma et al. 2026 ) .
Method
Put_block
Play_jenga
Open_jar
Close_box
Open_box
Pickup_cup
Push_block
VoxPoser ( Huang et al. 2023 )
70.70±2.31
0.00±0.00
0.00±0.00
0.00±0.00
0.00±0.00
26.70±14.00
25.33 ±8.33
CAP ( Liang et al. 2023 )
84.00±16.00
0.00±0.00
0.00±0.00
0.00±0.00
0.00±0.00
14.67±4.62
8.00±4.00
Scaling-up ( Ha et al. 2023 )
77.33±6.11
0.00±0.00
78.67±11.55
0.00±0.00
0.00±0.00
9.33±2.26
5.33±6.11
HiWE (Ours)
93.33 ±4.16
82.00 ±10.39
84.00 ±6.00
48.67 ±12.06
34.67 ±16.77
88.67 ±3.06
23.33±10.07
HiWE w/o FT
75.33±7.57
60.67±9.45
71.33±10.07
31.33±9.02
8.67±3.06
74.67±6.11
14.67±11.02
Method
Take_umbrella
Sort_mustard
Open_wine
Lamp_on
Put_knife
Pick_&_lift
Insert_block
Table 2: Task-averaged success rate % for zero-shot evaluation. HiWE outperformed other baselines in 10 out of 14 simulation tasks from RLBench ( James et al. 2020 ) . Each task was evaluated over 3 seeds to obtain the task-averaged success rate and standard deviations.
Figure 3: Illustrative manipulation sequences. The displayed stages connect object localization, a spatial motion plan, and robot execution for scenes involving several interactions.
Qwen-VL
LLaVA-NeXT
SpaceLLaVA
GPT-4o
PointVLM
24.1±0.9
20.0±0.9
21.3±0.9
15.3±1.3
52.1±1.2
Table 3: Quantitative comparisons on object reference (RoboRefIt). The metric is percentage of predicted points within the target mask.
No VQA
No LVIS
No Pixel
No Sim
No Robo
All
42.5±3.7
25.8±2.1
32.6±5.3
46.6±3.2
47.8±3.3
52.1±1.2
Table 4: Ablation on the data composition. Results on RoboRefIt show that best results are achieved when all of the data sources are combined during instruction-tuning.
Method
Move Spray Bottle
Open Drawer
Open Jar
Sort Object
CAP
6.67
0.00
36.67
70.00
RoboPoint
0.00
0.00
20.00
63.33
HiWE
60.00
33.33
56.67
80.00
Table 5: Zero-shot success rates (%) in real-world manipulation. HiWE consistently outperforms Code as Policies (CAP) ( Liang et al. 2023 ) and RoboPoint across four representative tasks.
Figure 4: The multi-view robustness of PointVLM. This assists 3DLLM in reasoning about the direction for pulling out the block.
Method
Take umbrella
Put block
3DLLM-2D
2.00±2.00
19.33±5.03
3DLLM-1point
26.67±3.06
82.00±4.00
3DLLM w/o obstacle
23.33±7.02
93.33±4.16
3DLLM
64.67±15.01
93.33±4.16
Table 6: Ablation Study on 3DLLM for Trajectory Planning