Organizations: Institute of Automation, Chinese Academy of Sciences, Beijing, China · School of Future Technology, University of Chinese Academy of Sciences · State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, CAS · The University of Hong Kong, Hong Kong SAR, China · Department of Electronic Engineering, Tsinghua University, China
Robot demonstration generation requires a system to identify where an interaction should occur, plan a feasible motion, and execute the required contact. HiWE connects these decisions through a point-based interface between visual grounding and language-based planning. PointVLM is instruction-tuned to associate task-relevant objects with image coordinates using a mixture of point annotations, segmentation-derived samples, robot observations, and visual question answering data. Depth measurements lift these predictions into a semantic 3D representation. A language planner, 3DLLM, uses this representation to specify end-effector waypoints and gripper commands, while a hybrid grasping module resolves local grasp poses. The evaluation covers 14 simulated manipulation tasks and four physical-robot tasks, together with ablations of the visual training data, spatial inputs, and grasp selection. Here, zero-shot execution refers to deployment without task-specific demonstration training; the visual model uses existing robot data during fine-tuning. This paper describes the original point-based formulation of the framework; its relationship to the subsequent GeneralVLA extension is detailed in the introduction.
Figures & tables
Figure 1: Interfaces for robot manipulation. HiWE exposes task-relevant points and planned trajectories between perception and execution. This separation allows the visual component to be trained on annotations that do not require deployment-specific action trajectories.
Component
HiWE
GeneralVLA
Perception
PointVLM coordinate prediction
ASM segmentation and refinement
Planning
Current task and 3D scene points
3D scene planning with KnowledgeBank
Experience
No persistent retrieval in 3DLLM
Retrieval, construction, consolidation
Execution
HGM grasp selection
HGM grasp selection
Table 1: Scope of the original system and its subsequent extension. Shared components are not independent evidence of a new architecture.
Figure 2: HiWE data interfaces. PointVLM associates task-relevant objects with image points. Depth supplies their 3D coordinates for 3DLLM, which outputs waypoints and gripper commands. HGM provides local grasp poses for the execution stage. The overall decomposition and grasping component are also described in GeneralVLA ( Ma et al. 2026 ) .
Method
Put_block
Play_jenga
Open_jar
Close_box
Open_box
Pickup_cup
Push_block
VoxPoser ( Huang et al. 2023 )
70.70±2.31
0.00±0.00
0.00±0.00
0.00±0.00
0.00±0.00
26.70±14.00
25.33 ±8.33
CAP ( Liang et al. 2023 )
84.00±16.00
0.00±0.00
0.00±0.00
0.00±0.00
0.00±0.00
14.67±4.62
8.00±4.00
Scaling-up ( Ha et al. 2023 )
77.33±6.11
0.00±0.00
78.67±11.55
0.00±0.00
0.00±0.00
9.33±2.26
5.33±6.11
HiWE (Ours)
93.33 ±4.16
82.00 ±10.39
84.00 ±6.00
48.67 ±12.06
34.67 ±16.77
88.67 ±3.06
23.33±10.07
HiWE w/o FT
75.33±7.57
60.67±9.45
71.33±10.07
31.33±9.02
8.67±3.06
74.67±6.11
14.67±11.02
Method
Take_umbrella
Sort_mustard
Open_wine
Lamp_on
Put_knife
Pick_&_lift
Insert_block
Table 2: Task-averaged success rate % for zero-shot evaluation. HiWE outperformed other baselines in 10 out of 14 simulation tasks from RLBench ( James et al. 2020 ) . Each task was evaluated over 3 seeds to obtain the task-averaged success rate and standard deviations.
Figure 3: Illustrative manipulation sequences. The displayed stages connect object localization, a spatial motion plan, and robot execution for scenes involving several interactions.
Qwen-VL
LLaVA-NeXT
SpaceLLaVA
GPT-4o
PointVLM
24.1±0.9
20.0±0.9
21.3±0.9
15.3±1.3
52.1±1.2
Table 3: Quantitative comparisons on object reference (RoboRefIt). The metric is percentage of predicted points within the target mask.
No VQA
No LVIS
No Pixel
No Sim
No Robo
All
42.5±3.7
25.8±2.1
32.6±5.3
46.6±3.2
47.8±3.3
52.1±1.2
Table 4: Ablation on the data composition. Results on RoboRefIt show that best results are achieved when all of the data sources are combined during instruction-tuning.
Method
Move Spray Bottle
Open Drawer
Open Jar
Sort Object
CAP
6.67
0.00
36.67
70.00
RoboPoint
0.00
0.00
20.00
63.33
HiWE
60.00
33.33
56.67
80.00
Table 5: Zero-shot success rates (%) in real-world manipulation. HiWE consistently outperforms Code as Policies (CAP) ( Liang et al. 2023 ) and RoboPoint across four representative tasks.
Figure 4: The multi-view robustness of PointVLM. This assists 3DLLM in reasoning about the direction for pulling out the block.
Method
Take umbrella
Put block
3DLLM-2D
2.00±2.00
19.33±5.03
3DLLM-1point
26.67±3.06
82.00±4.00
3DLLM w/o obstacle
23.33±7.02
93.33±4.16
3DLLM
64.67±15.01
93.33±4.16
Table 6: Ablation Study on 3DLLM for Trajectory Planning
World models such as DINO-WM and LeWM specify the goal with an image, which is difficult to obtain in advance for novel tasks. We present the Grounded World Model (GWM), a latent world model that enables zero-shot planning in the real world from language goals alone. Given a candidate action sequence and the current observation, GWM predicts the future in the visual space of a pretrained video-language embedding model. The frozen readout of this embedding model maps this imagined future and the task description into the same embedding space, where their negative cosine similarity serves as the planning cost. Training GWM requires only offline and task-agnostic video-action pairs and no language labels. In simulated experiments on WISER, planning with GWM, which executes the candidate action of lowest cost, solves 87% of 288 tasks with unseen instructions and visual signals, while ten fine-tuned VLAs average 22%. We then scale GWM up with real robot data, and use it for zero-shot planning in realistic simulation and real scenes. In the IsaacSim evaluation, planning with GWM completes all 70 trials across 14 tasks that require reasoning over referring expressions, matching a modular planner grounded by a frontier VLM, while pi0.5 reaches 37/70. Deployed on a real Franka, the same stack completes 55/60 separately evaluated pick-and-place sub-tasks, comparable to the modular planner's 52/60, with the full system running locally on a single consumer GPU. Project website: https://quanyili.github.io/gwm-wiser/.
We present ZeroDex, a zero-shot framework for long-horizon dexterous manipulation that grounds language instructions into executable 3D task plans from calibrated multi-view RGB images. Rather than training an end-to-end policy, our system uses a vision-language model (VLM) to produce reference-frame task grounding and primitive-level 2D keypoints, then lifts them into 3D via multi-view fusion. This lifting combines triangulation of view-wise VLM groundings with reference-view ray voting, which searches along a semantic camera ray for geometrically consistent candidates across neighboring views. The resulting 3D keypoints support both pick-and-place and tool-use: for tool-use, we retrieve an object-centric atomic action corresponding to the inferred skill category and align its stored 6D tool trajectory to the scene; for dexterous execution, we expand the lifted grasp keypoint into a task-conditioned grasp affordance region and generate feasible grasp-motion pairs with an arm-hand motion generator. Real-world experiments show improved 3D grounding accuracy and execution reliability over single-view RGB-D grounding and fine-tuned VLA baselines. We further demonstrate long-horizon manipulation through closed-loop status verification and replan, enabling zero-shot execution on unseen objects and tool-use tasks in novel scenes.
Intermediate representations are key to bridging the modality gap between generalizable manipulation policies and large-scale pretrained vision-language models (VLMs). Among these, trajectory-based representations compactly represent motion-relevant cues, yet most existing approaches predict trajectories in 2D image space, resulting in intrinsic 3D ambiguity. Moreover, using 2D trajectories with depth still leaves the free-space waypoints ambiguous, limiting reliable 3D reasoning. To address this, we propose predicting 3D consistent waypoints (3DWay) from multi-view images. By reformulating 3D waypoints prediction as generating multi-view consistent 2D waypoints followed by geometric triangulation, we enable explicit 3D motion specification while preserving the strong priors of pretrained VLMs. The predicted waypoints can guide existing VLA models for better generalization or be directly executed on simple tasks. Extensive experiments show that 3DWay substantially improves 3D spatial grounding and vision-language reasoning, demonstrating strong potential for generalizable robot manipulation. Codes will be released at https://github.com/ziqin-h/3DWay.
Ziqin Huang, Yingyue Li, Chenyangguang Zhang +6
Tsinghua University · ETH Zürich · University of California, Berkeley