Pretrained robot manipulation policies such as vision-language-action models (VLAs) or world-action models (WAMs) leave interaction-relevant metric geometry implicit. Recent breakthroughs in spatial reconstruction can supply the necessary geometry reliably, but their features describe local shape without stating where it lies with respect to the robot. How best to deliver these features to a pretrained policy remains unresolved. We propose Spatial Grafting, a versatile, lightweight spatial module that binds frozen reconstruction features to metric, robot-relative geometry. Spatial Grafting constructs metric-grounded spatial tokens and injects them into the flow-matching action expert through cross-attention, without modifying the host's perceptual pathway, so the host retains the full benefit of its pretraining. We evaluate it more broadly than any geometry-aware policy we compare against: one graft architecture, with no per-host redesign, on two VLAs and two WAMs, across four simulation benchmarks that span short-horizon manipulation, visual robustness, clutter and long-horizon mobile manipulation, and on three real-robot platforms with single- and dual-arm configurations. On RoboTwin 2.0, a dual-arm manipulation benchmark, the graft improves every host across VLAs and WAMs. Grafted π0.5 gains 11.3% and 15.6% on clean and randomized scenes, reaching 94.0% and 92.4%, above the strongest published 3D-conditioned policy, WAM4D (93.8% and 89.9%). The margin widens as the horizon lengthens: on tasks from BEHAVIOR-1K, a dual-arm mobile manipulation challenge scored by average task progress, it surpasses the 2025 challenge winner on five of six tasks,by up to 0.47 Q-score, and exceeds a map-conditioned spatial policy on average across the three tasks both report.
Figures & tables
Fig. 1: Overview of Spatial Grafting . A frozen geometry foundation model extracts multi-level features from each RGB view. Depth, calibrated rays, and end-effector anchors bind those features to metric, robot-relative 3D locations, forming a spatial bank per view. The magnified graft layer acts only on the flow-matching action expert: action tokens first cross-attend to the primary-view bank; the updated states then query the auxiliary-view banks. A single bias-free linear projection fuses the auxiliary attention outputs into the action update. This layer repeats in selected late action blocks, leaving the host’s vision-language or video pathway ungrafted.
Host
Grafted blocks
Bank builder
Injection
Bridges
Total (share)
π0.5
6
46.0M
100.7M
–
146.8M ( 4.4% of 3.3B)
X-VLA
6
46.0M
100.7M
–
146.8M
π0.5 , VGGT- Ω
6
48.1M
100.7M
–
148.9M
Fast-WAM
10
46.0M
168.0M
–
214.0M
LingBot-VA
10
46.0M
168.0M
62.9M
276.9M ( 5.4% of 5.09B)
TABLE I: Trainable parameters added by the graft on RoboTwin. Counts are measured by building the modules. The share is relative to the host’s own parameter count.
Model
LIBERO Avg.
RoboTwin Clean
RoboTwin Rand.
GLaD
94.1
NR
NR
GeoVLA
97.7
NR
NR
SpatialVLA
78.1
NR
NR
Spatial Forcing
98.5
66.2
58.9
WAM4D
NR
93.8
89.9
π0.5→π0.5graft
96.9 →
99.6 +2.7
82.7 →
94.0 +11.3
76.8 →
92.4 +15.6
TABLE II: LIBERO and RoboTwin SR (%). Above the line: published 3D-conditioned policies, quoted from their own papers, except Spatial Forcing on RoboTwin, which we trained ourselves (Section VII ). Below the line: our four hosts. LIBERO bases are quoted from each host’s published results; the π0.5 entry is from the openpi release [ 41 ] , as its paper reports no LIBERO result. On RoboTwin, the π0.5 base is cited from [ 14 ] and the X-VLA base from [ 42 ] , both trained on the same RoboTwin demonstrations as our grafts, so the pairs are directly comparable; the Fast-WAM and LingBot-VA bases are our evaluations of their public checkpoints with our seeds and pipeline. Every grafted entry is our run. Green/red: absolute change from the host’s own base. NR: not reported in the cited paper.
Task
SERF
π0.5-Comet
π0.5-RLC→
π0.5-RLCgraft
Assembling gift baskets
0.525
0.312
0.281 →
0.632 +0.351
Collecting children’s toys
0.635
0.583
0.500 →
0.673 +0.173
Putting shoes on rack
0.601
0.470
0.720 →
0.590 −0.130
Putting up Christmas decorations
NR
0.356
0.478 →
0.510 +0.032
Clean boxing gloves
NR
0.000
0.200 →
0.600 +0.400
Cleaning up plates and food
NR
0.043
0.243 →
0.714 +0.471
TABLE III: B1K Q-score on six tasks. Q-score is average task progress. π0.5-RLC and π0.5-Comet are the #1 and #2 entries of the 2025 challenge [ 8 ] and SERF [ 25 ] is quoted on the three tasks it reports; all three are taken from their published scores. The grafted policy is evaluated on the 20 public evaluation instances of each task; official challenge scoring uses only 10 of those 20. Green/red: change from the host’s own base.
Fig. 2: Grafted success (top) against base failure (bottom), one column per benchmark. Each column is the task discussed in the corresponding paragraph. R1Pro: the grafted policy estimates plate height and releases the orange with clearance, where the base drops it short. Piper: plate–rack alignment avoids the collision the base drives into. B1K: the base shuts its jaws beside the candy cane’s shaft; the grafted policy closes on it. RoboTwin: rack height is estimated well enough to clear the hook. RoboPRO: a lower, better-placed grasp seats the plate in the sink instead of catching its rim.
Model
Clean (Easy/Hard)
Clutter (Easy/Hard)
π0.5→π0.5graft
70.3 / 65.7 →
78.8 / 73.7 +8.5/+8.0
60.9 / 47.8 →
69.8 / 49.4 +8.9/+1.6
X-VLA → X-VLA graft
49.5 / 46.7 →
60.9 / 49.9 +11.4/+3.2
39.7 / 27.6 →
45.3 / 35.3 +5.6/+7.7
Spatial Grafting : base → graft
TABLE IV: RoboPRO success (%). Each cell reads Easy/Hard. Easy is RoboPRO’s success rate (SR); Hard is its stricter harsh success rate (HSR), scored on the same episodes. Bases are reported by [ 7 ] and grafted scores are ours. Green: change from the host’s own base.
Platform
Task
SR
Trials
UR5e
Pick up marker and place in pen cup (thin objects, constrained placement)
30 →
80 +50
10
Piper
Clean up table (transparent objects, long horizon)
20 →
70 +50
10
Open tea can (bimanual dexterous manipulation)
10 →
40 +30
10
Place plate on rack (thin objects, constrained placement)
70 →
90 +20
10
R1Pro
Pick up tangerine and place on plate (thin objects, constrained placement)
50 →
90 +40
10
Clean up table (deformable object, long horizon)
50 →
80 +30
10
TABLE V: Real-robot evaluation (SR, %). Matched base/graft comparisons on three platforms, with 10 trials per task. A trial is considered successful only if all sub-tasks are completed. Green: change from the base.
Fig. 3: Ablations on RoboTwin with π0.5 (SR %). (a) Reference, ungrafted base, and bank interventions on a fixed checkpoint. (b) One design choice retrained at a time. Equal-width Clean (green) and Randomized (terracotta) bars overlap at each configuration and share a 25% baseline.
Fig. 4: Success and failure modes across platforms. Left and right show successful and failed rollouts. For R1Pro fruit placement, improved plate-height estimation leaves clearance above the rim, while the failure approaches too low. For Piper plate placement, better plate–rack alignment avoids collision; the two panels use left- and right-wrist fisheye cameras, respectively. In B1K candy pickup and RoboPRO plate pickup, target depth supports a lower grasp pose. In RoboTwin mug hanging, estimating the rack height positions the mug at the support. The added Piper tea-can opening and UR5e marker-grasp panels are cropped and AI-enhanced for visual clarity; generative deblurring may reconstruct fine details, so these panels are illustrative rather than unaltered experimental frames.
Method
Status
Geometry signal
Route into the model
Test-time inputs
Evaluation
Hosts: VLAs and WAMs
π0.5
CoRL’25
None
VLM trunk co-trained with a flow-matching action expert
RGB
Real
X-VLA
ICLR’26
None
Embodiment-specific soft prompts into a flow-matching transformer
RGB
LIBERO; SimplerEnv; CALVIN; VLABench; RoboTwin; Real
Fast-WAM
arXiv’26
None
Video and action experts in a shared mixture of transformers; no video generation at inference
RGB
LIBERO; RoboTwin; Real
LingBot-VA
RSS’26
None
Autoregressive video/action tokens in a mixture of transformers
RGB
LIBERO; RoboTwin; Real
Spatial action models compared in Section V
TABLE VI: Host policies and geometry-aware policies. Rows follow Section II . Status is the venue as of submission; “Real” denotes real-robot experiments; test-time inputs are what the deployed policy consumes beyond proprioception. The four hosts carry no geometry module. Every geometry-aware method reports at most three evaluation settings; Spatial Grafting reports five, on four hosts.
Method
Spat.
Obj.
Goal
Long
Avg.
π0
96.8
98.8
95.8
85.2
94.1
OpenVLA-OFT
97.6
98.4
97.9
94.5
97.1
GLaD
95.0
97.4
94.4
89.4
94.1
GeoVLA
98.4
99.0
96.6
96.6
97.7
Spatial Forcing
99.4
99.6
98.8
96.0
98.5
SpatialVLA
88.2
89.9
78.6
55.5
78.1
TABLE VII: Detailed LIBERO SR (%). Results across the four standard suites. Values for external methods are reported by their respective works; the π0.5 base is from the openpi release [ 41 ] . Host averages match Table II ; dashes indicate unavailable per-suite results, not zero success. Blocks group other baselines, existing 3D methods, and our hosts. Green/red: change from the host’s own base.
Task
π0.5graft
X-VLA graft
Fast-WAM
Fast- WAM graft
LingBot-VA
LingBot- VA graft
π0.5graft (VGGT- Ω )
Adjust bottle
100/100
100/95
100/95
100/100
100/100
95/95
100/100
Beat block hammer
100/100
90/85
100/100
100/95
100/100
100/100
100/95
Blocks ranking RGB
100/95
100/85
95/85
100/90
75/90
100/100
100/100
Blocks ranking size
85/70
75/85
60/65
65/85
85/50
95/90
85/75
Click alarmclock
95/100
10/30
100/100
100/100
100/100
100/100
95/100
Click bell
100/100
70/90
100/100
100/100
100/100
100/100
95/100
TABLE VIII: Per-task RoboTwin SR (%), Clean/Randomized. Each cell reads Clean/Randomized over exactly 20 seeds per task; every model and evaluation uses the same seeds. DA3 is the spatial backbone unless marked VGGT- Ω ; Fast-WAM and LingBot-VA without a superscript are the ungrafted public checkpoints. The average is the mean of the 50 per-task rates.
π0.5graft
X-VLA graft
Task
Clean
Clutter
Clean
Clutter
Chain apple bin bowl rack spoon sink
15/15
10/10
0/0
0/0
Chain apple sink plate bread board
95/85
100/90
0/0
0/0
Chain bowl rack apple sink
0/0
0/0
0/0
0/0
Chain heat hamburger
45/45
85/75
20/20
0/0
Chain serve hamburger
45/45
0/0
0/0
0/0
TABLE IX: Per-task RoboPRO success (%), Easy/Hard. Each cell reads Easy (SR)/Hard (HSR) over exactly 20 seeds per task for all 80 tasks in each configuration, the same seeds for every model; averages are the mean of the 80 per-task rates and match Table IV .
Robotic manipulation requires reasoning about future spatial-temporal interactions and geometric constraints, yet existing Vision-Language-Action (VLA) policies often leave predictive representation weakly coupled with action execution, causing failures in tasks requiring precise spatial-temporal coordination. We propose STARRY, a world-model-enhanced action-generation policy that aligns spatial-temporal prediction and action generation by jointly denoising future spatial-temporal latents and actions through a unified diffusion process. To bridge 2D visual tokens and 3D metric control, STARRY introduces Geometry-Aware Selective Attention Modulation (GASAM), which converts predicted depth and end-effector geometry into token-aligned weights for selective action-attention modulation. On RoboTwin 2.0, STARRY achieves 93.82% / 93.30% average success under Clean and Randomized settings across 50 bimanual tasks. Real-world experiments show that STARRY improves average success from 42.5% to 70.8% compared with π0.5. These results demonstrate the effectiveness of action-centric spatial-temporal world modeling for spatially and temporally demanding robotic manipulation.
Yuxuan Tian, Yurun Jin, Bin Yu +5
Beijing Institute of Technology · Zhongguancun Academy · Zhongguancun Institute of Artificial Intelligence +4
Robotic manipulation policies are advancing rapidly with increasing reliance on vision-language models for end-to-end decision making. However, reliable deployment remains challenging because many policies lack explicit mechanisms for predicting task outcomes and evaluating whether generated actions will achieve desired final states, causing execution errors to accumulate during long-horizon manipulation. We present Robot-GST, a geometry-aware spatio-temporal behaviour representation and evaluation framework that constructs a Gaussian-SAM robotic environment for real-to-sim policy verification and improves the reliability of real-world manipulation deployment. Our approach constructs a high-fidelity robotic environment from RGB-D observations using 3D Gaussian Splatting and SAM3D, enabling ``simulation and evaluation before acting''. It integrates visual observations and language instructions with spatio-temporal reasoning for long-horizon task planning using large vision-language models. To bridge high-level planning and real-world execution, we introduce Gaussian-aware final-state estimation through geometric sampling and state-based trajectory planning. Before execution, candidate action sequences are simulated and evaluated in the Gaussian-SAM environment to filter infeasible behaviours. We validate our approach on representative manipulation tasks involving rigid, soft, and deformable objects, including cube placing, toy packing, and duck rearrangement, demonstrating that geometry-aware spatio-temporal reasoning and state-aware execution improve manipulation reliability across different object categories. Our results suggest that combining geometry-aware reconstruction with high-quality rendering and simulation provides a scalable approach for evaluating robotic manipulation behaviours. Website: https://robot-gst.github.io
Vision-Language-Action (VLA) models leverage large-scale vision-language pretraining for flexible robot manipulation, yet at test time they remain brittle to changed object positions and to familiar scenes paired with different instructions. A growing family of methods addresses this brittleness by supplying the policy with grounding signals, such as 2D pixel coordinates for object localization and placement. However, we find that how the grounding signal is represented and injected matters more than the signal itself. In this work, we propose a lightweight module that represents the grounding signal in 3D and injects the resulting embedding directly into the action head. The module is a two-layer MLP and requires no changes to the VLA backbone or pretraining pipeline, yet it yields substantially larger gains than language- or visual-prompting alternatives. On LIBERO-PRO, our method improves the average success rate of GR00T-N1.6 from 31.2 to 77.5 under task perturbation and from 28.1 to 60.2 under position perturbation. Comparable gains are also achieved for π0.5, demonstrating that the mechanism is backbone-agnostic across VLAs with diffusion-based action heads. We further validate the practical applicability with real-world experiments. Together, these results support our central finding: lifting adequate 2D grounding into 3D and injecting it into the action head enables spatial and instance-level task generalization in VLAs.
Shiang-Feng Tsai, Jin-Cheng Jhang, Yen-Ling Tai +5
National Tsing Hua University · National Yang Ming Chiao Tung University