cs.ROSep 23, 2025

Agentic Scene Policies

Authors: Sacha Morin, Kumaraditya Gupta, Mahtab Sandhu, Charlie Gauthier, Francesco Argenziano, Kirsty Ellis, Liam Paull

Organizations: Université de Montréal · Mila - Quebec AI Institute · Sapienza University of Rome

Abstract

Designing or learning robot policies that generalize zero-shot across a range of language instructions and objects is a core problem in robotics. Vision-Language-Action models (VLAs) learn such policies end-to-end by repurposing existing Vision-Language Models (VLMs), but generalization to new instructions and objects remains challenging. An alternative is to implement a modular policy by leveraging an explicit VLM-based 3D scene representation and motion planning. While modular policies show strong zero-shot potential, they typically retrieve objects based on semantics without explicit spatial reasoning, severely restricting their overall grounding capabilities. They also interact with objects using basic grasping and navigation skills. In this work, we address these limitations by unifying grounding capabilities and robot skills in a single agentic action space through a scene-agent tool interface. By leveraging part-level affordances, our skills generalize across diverse objects and enable zero-shot interactions such as unplugging chargers and opening drawers. We name the resulting framework Agentic Scene Policies (ASP). Through extensive real-world experiments, we show how ASP consistently outperforms leading VLAs in the zero-shot setting. We also demonstrate the extensibility of our framework by introducing a mobile version of ASP to tackle room-level queries. See our project page (https://montrealrobotics.ca/agentic-scene-policies.github.io/) for more results.

Explore similar work

CardsList
  1. Decoupling the Declarative from the Procedural in Vision-Language-Action Models

    Jun 19, 2026Nikolaos Tsagkas, Andreas Sochopoulos, Chris Xiaoxuan Lu +2Diffusion-Based Vision-Language-ActionsLarge-Scale Robot Demonstration Datasets

  2. In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use

    Aug 6, 2026Jiarui Yang, Wen Huang, Jiale Zhang +2Behavior Cloning