cs.ROJun 12, 2025

Gondola: Grounded Vision Language Planning for Robotic Manipulation

Authors: Shizhe Chen, Ricardo Garcia, Paul Pacaud, Cordelia Schmid

Organizations: Inria, École normale supérieure, CNRS, PSL Research University

Abstract

Vision-language-action (VLA) models have shown promising progress in robotic manipulation. However, directly mapping visual observations and language instructions to low-level actions often results in limited interpretability and weak robustness in complex, long-horizon tasks. To address these challenges, we employ a modular manipulation framework that separates high-level planning from low-level control. At its core is Gondola, a grounded vision-language planning model that generates structured plans with explicit pixel-level object grounding before action execution. Given multi-view observations and planning history, Gondola predicts the next-step plan as interleaved textual instructions and multi-view segmentation masks corresponding to target objects and goal locations. To train Gondola, we construct synthetic datasets that provide explicit supervision for short-horizon grounded planning, multi-view referring expression, and long-horizon compositional reasoning. By coupling grounded plan generation with a 3D-based execution policy, our framework achieves state-of-the-art performance on the challenging GemBench benchmark. The system further demonstrates promising transfer to real robots. Ablation studies confirm that pixel-level grounding and the proposed planning-oriented supervision are critical for effective high-level reasoning. Project webpage: https://cshizhe.github.io/projects/robot_gondola.html

Figures & tables

Explore similar work

CardsList
  1. ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation

    Jun 16, 2026Tianyi Lu, Hui Zhang, Zijie Diao +8Diffusion-Based Vision-Language-ActionsRobotic Manipulation

  2. GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations

    May 21, 2026Wenxuan Guo, Ziyuan Li, Meng Zhang +7Robotic Manipulation

  3. Grounded Action Model: 3D Grounding as a Foundation for Robotics

    Sep 20, 2026Gehao Zhang, Weikai Huang, Shailesh Shailesh +33D Visual Grounding