Vision-language-action (VLA) models have shown promising progress in robotic manipulation. However, directly mapping visual observations and language instructions to low-level actions often results in limited interpretability and weak robustness in complex, long-horizon tasks. To address these challenges, we employ a modular manipulation framework that separates high-level planning from low-level control. At its core is Gondola, a grounded vision-language planning model that generates structured plans with explicit pixel-level object grounding before action execution. Given multi-view observations and planning history, Gondola predicts the next-step plan as interleaved textual instructions and multi-view segmentation masks corresponding to target objects and goal locations. To train Gondola, we construct synthetic datasets that provide explicit supervision for short-horizon grounded planning, multi-view referring expression, and long-horizon compositional reasoning. By coupling grounded plan generation with a 3D-based execution policy, our framework achieves state-of-the-art performance on the challenging GemBench benchmark. The system further demonstrates promising transfer to real robots. Ablation studies confirm that pixel-level grounding and the proposed planning-oriented supervision are critical for effective high-level reasoning. Project webpage: https://cshizhe.github.io/projects/robot_gondola.html
Figures & tables
Fig. 1 : Gondola simultaneously produces an interpretable, mask-grounded structured plan that localizes task-relevant objects and placements across views.
Fig. 2 : Left: Gondola model architecture, consisting of a shared visual encoder for multi-view images, an LLM to generate action and object names along with segmentation tokens, and SAM2 to decode masks. Right: Integrating Gondola with a 3D-based motion planning policy for closed-loop task execution.
Fig. 3 : Constructed datasets: (1) robot grounded planning, (2) multi-view referring expressions for improved object grounding, and (3) pseudo long-horizon tasks for enhanced long-horizon reasoning.
L1
L2
L3
L4
Grd type
Multi- view
Hist- ory
Act
Obj
Grd
Act
Obj
Grd
Act
Obj
Grd
Act
Obj
Grd
Box
✓
✓
95.1
93.7
62.8
97.4
89.0
58.7
69.3
53.8
46.2
70.0
36.8
16.6
Mask
×
×
98.0
98.2
87.8
95.1
89.5
79.8
79.3
76.3
62.7
77.2
40.0
37.4
Mask
✓
×
100
100
88.6
98.0
91.3
81.2
85.6
75.6
61.4
89.9
50.1
46.5
Mask
✓
✓
100
100
87.9
99.0
93.3
79.2
88.6
83.9
66.4
79.4
44.9
40.0
TABLE I : Performance on grounded planning evaluation. We measure the action (Act) and object (Obj) name prediction accuracy and grounding performance (Grd) on the four levels of GemBench validation split. All the models are fine-tuned on the robot grounded planning dataset.
Finetuning Data
L1
L2
L3
L4
Plan
RefExp
Long
Act
Obj
Grd
Act
Obj
Grd
Act
Obj
Grd
Act
Obj
Grd
✓
×
×
100
100
87.9
99.0
93.3
79.2
88.6
83.9
66.4
79.4
44.9
40.0
✓
✓
×
100
100
89.1
99.0
95.1
84.2
92.1
88.2
73.3
72.1
42.3
41.7
✓
✓
✓
100
100
89.5
99.7
95.3
85.2
88.5
82.2
73.8
93.9
51.2
53.8
TABLE II: Performance on grounded planning evaluation. All the models use multi-view and history plans for mask generation, but are fine-tuned on different composition of datasets: robot grounded planning (Plan), multi-view referring expression (RefExp), and pseudo long-horizon tasks (Long).
Fig. 4 : Averaged pixel-level grounding performance (mIoU) in GemBench as a function of training data size.
Method
FT
GT obj.
L1
L2
L3
L4
RoboPoint [ 21 ]
×
✓
23.0
22.0
18.6
18.2
LISA [ 22 ]
×
✓
25.7
24.0
19.5
24.8
LISA [ 22 ]
✓
✓
62.3
56.7
37.8
43.7
Gondola (Ours)
✓
×
89.5
85.2
73.8
53.8
TABLE III : Offline object grounding comparison on GemBench. ‘FT’ denotes fine-tuning on our synthetic datasets, and ‘GT obj.’ denotes using ground-truth object queries as prior grounding VLMs cannot predict task plans.
Method
3D
L1
L2
L3
L4
w/o LLM
Hiveformer [ 32 ]
✓
60.3 ±1.5
26.1 ±1.4
35.1 ±1.7
0.0 ±0.0
PolarNet [ 35 ]
✓
77.7 ±0.9
37.1 ±1.4
38.5 ±1.7
0.1 ±0.2
3D diffuser actor [ 38 ]
✓
91.9 ±0.8
43.4 ±2.8
37.0 ±2.2
0.0 ±0.0
RVT-2 [ 39 ]
✓
89.1 ±0.8
51.0 ±2.3
36.0 ±2.2
0.0 ±0.0
3D-LOTUS [ 14 ]
✓
94.3 ±1.4
49.9 ±2.2
38.1 ±1.1
0.3 ±0.3
w/ LLM
GR00T N1.5 [ 7 ]
×
20.9 ±1.2
6.8 ±1.6
26.5 ±1.6
0.0 ±0.0
TABLE IV: Success rate of task execution on four levels of GemBench testing split.
Act chunk
Unified
L1
L2
L3
L4
×
80.8 ±1.2
68.5 ±1.5
48.6 ±1.1
4.1 ±1.3
5
✓
87.3 ±1.9
74.8 ±1.8
52.4 ±2.1
19.0 ±1.0
1
✓
90.8 ±1.2
78.2 ±1.4
49.5 ±0.5
14.9 ±2.2
TABLE V: Ablation study on unified planning–grounding and action chunk size. Unified modeling improves performance, and replanning frequency affects the trade-off between fine-grained control and long-horizon consistency.
Fig. 5 : Real-robot experimental setup and models’ performance on real-robot tasks, evaluated by success rate (SR).