Spatial reasoning is fundamental to general robot intelligence, as it enables robots to complete long-horizon tasks involving multi-object interaction. We introduce Simify, a training-free, test-time framework that performs explicit spatial reasoning via massively parallel physics simulation. From a single RGB-D image of a scene, Simify reconstructs simulation-ready assets leveraging 3D generative models and vision-language models. Then given a task specified by a reward function (e.g., build the tallest tower), Simify launches thousands of parallel rollouts in simulation and performs an evolutionary search to optimize object arrangements, typically converging within seconds. We conduct quantitative experiments on real-robot hardware to demonstrate the ability of our framework to execute complex object rearrangement tasks end-to-end with previously unseen objects. Results show that our framework outperforms prior work on foundation models for spatial reasoning by effectively exploiting large-scale parallel simulation during inference, and also highlight the importance of complete and accurate geometry for successful sim-to-real transfer.
Figures & tables
Fig. 1: Zero-shot spatial and physical reasoning with Simify. From a single RGB-D view and task reward, the robot reconstructs the scene, searches for object arrangements in simulation, and then executes the rearrangement in the real world.
Fig. 2: Overview of Simify. (1) From a single RGB-D view, the system segments and completes objects, generates scaled meshes, estimates their poses, and builds a parallel physics simulation. (2) Evolutionary search jointly optimizes placement order and goal object poses, followed by noisy validation to select a robust, high-reward arrangement. (3) The robot tracks objects, samples grasps, and executes collision-free pick-and-place motions.
Fig. 3: Normalized task scores for Simify, LLM-GROP , and TSDF-EA . LLM-GROP uses a VLM for goal-pose reasoning, while TSDF-EA uses partial TSDF geometry. Simify consistently finds higher-scoring arrangements.
Fig. 4: Real-robot results over 10 trials per task. (A) Mean task progress. (B) Failure-stage distribution, where Stage 1 is planning and later stages are sequential object placements. Simify outperforms TSDF-EA and more often reaches later execution stages.
Method
Stacking
Cantilever
Dominoes
Mean
MeshGen-Rand
29.8
0.0
1.1
10.3
Simify-NoVal
62.6
45.6
83.8
64.0
Simify (ours)
82.2
85.9
94.3
87.5
TABLE I: Task score (%) results for an ablation study, exploring the design choices in our evolutionary algorithm.
Fig. 5: Failure breakdown for Simify over 30 real-world trials.
Fig. 7: Effect of simulation parallelism on task score. More environments improve performance on all tasks, showing that Simify benefits from additional test-time compute.
Chamfer Completeness ↓
Chamfer Accuracy ↓
F1-Score ↑
Normal Consistency ↑
Low Occ.
TSDF [ 42 ]
Bottle
0.015
0.003
0.687
0.752
Bowl
0.009
0.003
0.801
0.787
Box
0.026
0.002
0.596
0.791
Pitcher
0.022
0.002
0.642
0.728
Mean
0.018
0.002
0.682
0.765
Mesh Gen. (ours)
Bottle
0.004
0.004
0.939
0.917
TABLE II: Surface reconstruction quality of generated and partial TSDF meshes, measured against ground-truth YCB meshes.
Fig. 8: Qualitative comparison of partial TSDF, SAM3D, and Hunyuan3D-2 meshes at different octree resolutions. We use Hunyuan3D-2 at resolution 64 to balance reconstruction speed and accuracy; the modular pipeline also supports alternative image-to-3D models.
We present ZeroBot, a real2sim framework for learning a robot manipulation task from scratch in minutes under challenging conditions: zero human demonstrations, zero policy pre-training, and zero known object models. Given only a single view of an object and a goal pose for that object, ZeroBot uses image-to-3D generative models to obtain a complete object mesh, which is used in simulation for large-scale parallel reinforcement learning. To accelerate training, we introduce an action space which leverages the generated geometry and learned value function to sample states involving robot-object contact. When evaluated on real-world tasks including grasping, pushing, articulated object interaction, and multi-stage manipulation, ZeroBot achieves an 87% success rate with an average training time of 119 seconds. These results show the value of using image-to-3D models in a real2sim framework for rapid, autonomous robot learning.
Ivan Kapelyukh, Xiaohan Zhang, Stephen James +2
Imperial College London · Robotics and AI Institute
Humans can effortlessly perceive spatial layouts, form cognitive representations, reason about spatial relations, and translate such reasoning into actions in everyday 3D environments. Although recent vision-language models (VLMs) have shown promising performance on observation-conditioned spatial perception and reasoning tasks, it remains unclear whether they can build coherent spatial understanding, act upon it, and refine their actions through multi-turn feedback. To study this problem, we introduce \textbf{SpatialAct}, a simulator-grounded benchmark for probing \textit{action-conditioned spatial reasoning} in 3D scenes. Starting from the most challenging setting, Multi-turn Interactive Refinement, we further design its decomposed counterpart, Single-step Error Detection and Fix, together with five fundamental spatial ability tasks to diagnose the underlying causes of model failures. Experiments reveal a clear reasoning-to-action gap: current VLMs can perform well on isolated spatial reasoning tasks, but struggle to maintain coherent spatial beliefs and produce reliable actions during multi-turn feedback, substantially underperforming humans. These results suggest that current VLM agents still lack robust spatial state tracking under action-induced environment changes, even when low-level control is abstracted away.
Tianhui Liu, Jie Feng, Zhiheng Zheng +6
1The Hong Kong University of Science and Technology (Guangzhou) · 2Zhongguancun Academy · 3Tsinghua University +1
Real-world videos provide rich demonstrations of manipulation, but turning them into reusable robot skills requires visually aligned environments, executable physical interactions, and mechanisms for learning from experience. We introduce Real2Gym, an agentic Real2Sim2Real framework that turns human and robot demonstrations into interactive simulation gyms and brings skills acquired in simulation to physical robots. The Real2Sim module reconstructs editable scenes, aligns objects and cameras with the input, validates demonstrated or retargeted actions through native physics execution, and generates task-conditioned variations with action-feasibility checks. Within these environments, the agent generates executable code for manipulation stages, observes their outcomes, and distills successful attempts and failures into reusable task procedures, object-relative motions, and recovery strategies. Through a shared perception-and-control interface, these skills guide subsequent execution in simulation and on real robots, with motions adapted to current observations and no updates to the underlying model weights. Extensive evaluations demonstrate that Real2Gym enables high-fidelity simulation environment reconstruction, outperforming GPT-6 Astra Direct Mode by 16.7% in success rate with approximately 74.9% fewer policy-execution tokens across these environments, while exceeding it by 33.3% in physical robot execution success rate across four tasks on a real Franka robot.
Kerui Ren, Yingxiang Xu, Kaiwen Song +4
Shanghai Artificial Intelligence Laboratory · Shanghai Jiao Tong University · Zhejiang University +3