Organizations: College of Computer Science and Artificial Intelligence, Fudan University · Singapore Management University · Institute of Trustworthy Embodied AI, Fudan University
Frontier multimodal foundation models (e.g., GPT-6 Astra) have recently shown strong potential for direct robotic control, yet their performance on fine manipulation remains limited. We argue that an important source of failure is not necessarily insufficient policy capability, but insufficient spatial observability, where task-critical spatial relationships may be poorly revealed by the existing physical camera setup. We introduce SpatialHarness, a test-time embodied harness that provides test-time spatial scaffolding for fine robotic manipulation without policy fine-tuning or changes to the physical sensing setup. SpatialHarness maintains an online simulated scene synchronized with real-world execution, identifies task-critical spatial relationships, and renders complementary virtual views that expose them to a frozen multimodal policy. To keep the simulated scene aligned during interaction, we develop interaction-aware scene synchronization that distinguishes static, held, and transition modes. We evaluate SpatialHarness on four real-robot manipulation tasks spanning precise geometric alignment, object-relative placement, and articulated-object interaction. Using the same frozen GPT-6 Astra policy, SpatialHarness substantially improves task success, including from 26.7% to 66.7% on plug insertion and from 0% to 100% on Tower of Hanoi. These results indicate that improving spatial observability at test time can unlock fine-manipulation capabilities already present in strong multimodal foundation models. Project website: https://emilia113.github.io/SpatialHarness/.
Figures & tables
Figure 1 : The top two rows show real and auxiliary virtual views, respectively, for the four tasks shown; cyan marks task-relevant regions. Below, baseline (green) and SpatialHarness (yellow) success rates are compared using the same GPT Astra Low-based manipulation policy and real camera configuration. Values above the bars show gains over the baseline in percentage points (pp).
Figure 2 : Online scene synchronization by interaction mode. (a) Measured robot motion advances the physics simulation, updating the scene state and rendering supplementary views. (b) Stationary object poses are retained or visually aligned according to movability. For continuously grasped objects, poses are predicted through physics simulation and validated; rejected predictions are replaced with visual estimates. When a grasp is established or released, the same measured motion segment is replayed from candidate pre-interaction poses to select the current state by interaction consistency and visual matching. (c) Real oblique and corresponding virtual side and top views of Tower of Hanoi. Columns run chronologically from left to right, showing scene updates during real execution and hole–peg spatial relationships.
Figure 3 : Real robot tasks and experimental platform. Left: initial and final states of block stacking, plug insertion, Tower of Hanoi, and drawer opening followed by block placement. Right: the Franka Research 3 arm, camera positions, and Tower of Hanoi setup.
Figure 4 : Success rates on real robot tasks. The baseline and SpatialHarness use the same manipulation policy and control protocol, with 15 trials per task and method. The rightmost column gives the average across the four tasks. Parentheses give successes/total trials; arrows show gains over the baseline in percentage points (pp).
Figure 5 : Qualitative comparison of the baseline and SpatialHarness on four real-world manipulation tasks: (1) block stacking, (2) plug insertion, (3) Tower of Hanoi, and (4) block in drawer. Each row shows representative frames from task execution. Red circles highlight failure locations, while crosses and checkmarks indicate task failure and success, respectively.
Figure 6 : Success trials versus decision calls. SpatialHarness lies above and to the left of the baseline for each plotted task, combining higher success with fewer decision calls. Colors identify methods, marker shapes identify tasks, and stars denote task-average summaries.