cs.ROSep 28, 2026

Spatial Grafting: Grounding 3D Features for Flow-Matching Robot Policies

Authors: Dingsheng Liu, Yangzheng Wu, Mahboubeh Asadi, Zhiyuan Li, Jinbang Huang, Yixin Xiao, Tongtong Cao, Yingxue Zhang

Organizations: University of Toronto · Huawei Noah’s Ark Lab · Department of Foundation Model, 2012 Labs

Abstract

Pretrained robot manipulation policies such as vision-language-action models (VLAs) or world-action models (WAMs) leave interaction-relevant metric geometry implicit. Recent breakthroughs in spatial reconstruction can supply the necessary geometry reliably, but their features describe local shape without stating where it lies with respect to the robot. How best to deliver these features to a pretrained policy remains unresolved. We propose Spatial Grafting, a versatile, lightweight spatial module that binds frozen reconstruction features to metric, robot-relative geometry. Spatial Grafting constructs metric-grounded spatial tokens and injects them into the flow-matching action expert through cross-attention, without modifying the host's perceptual pathway, so the host retains the full benefit of its pretraining. We evaluate it more broadly than any geometry-aware policy we compare against: one graft architecture, with no per-host redesign, on two VLAs and two WAMs, across four simulation benchmarks that span short-horizon manipulation, visual robustness, clutter and long-horizon mobile manipulation, and on three real-robot platforms with single- and dual-arm configurations. On RoboTwin 2.0, a dual-arm manipulation benchmark, the graft improves every host across VLAs and WAMs. Grafted π0.5π_{0.5} gains 11.3% and 15.6% on clean and randomized scenes, reaching 94.0% and 92.4%, above the strongest published 3D-conditioned policy, WAM4D (93.8% and 89.9%). The margin widens as the horizon lengthens: on tasks from BEHAVIOR-1K, a dual-arm mobile manipulation challenge scored by average task progress, it surpasses the 2025 challenge winner on five of six tasks,by up to 0.47 Q-score, and exceeds a map-conditioned spatial policy on average across the three tasks both report.

Figures & tables

Explore similar work

CardsList
  1. STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation

    Apr 29, 2026Yuxuan Tian, Yurun Jin, Bin Yu +5Robotic ManipulationWorld Models

  2. Robot-GST: geometry-aware spatial-temporal robot policy representation and evaluation

    Sep 27, 2026Sichao Liu, Zekun Wang, Lixuan Tang +6Robotic Manipulation PoliciesRobot Systems

  3. Direct Action-Head Injection of A Grounded 3D Point Unlocks Spatial and Task Generalization

    Jun 26, 2026Shiang-Feng Tsai, Jin-Cheng Jhang, Yen-Ling Tai +5Robotic ManipulationRecent Vision-Language Models