A simulation of a real robot workspace must preserve task-relevant interactions, while policies developed in it must operate on observations available to the real robot. Yet scene reconstruction and policy development are often treated separately. We present Agentic Real-to-Sim-to-Real (Agentic RSR), a framework that links scene reconstruction, policy development, and real-robot execution through the same manipulation task. Given a workspace video, a task description, and a known robot model, an agent recovers metric scale, iteratively refines the scene using visual feedback, and checks task-relevant interactions in MuJoCo. A coding agent then develops an executable policy, progressing from privileged object poses to visual observations and randomized simulation. The policy can interleave multiple observations and actions within one invocation, while the agent uses execution feedback to continue, retry, or revise its approach. A shared task-level interface carries the policy and accumulated experience to the real robot, where fresh observations and safety checks guide execution. Across 18 reconstructed scenes involving two robots, the mean four-view Depth MAE against reference depth estimates is 0.1057 m, the mean Lab ΔE76 is 11.04, and the mean grayscale SSIM is 0.6990. In real-robot experiments, the aggregate task success rate reaches 80% of the simulation task success rate, indicating substantial retention of simulated performance on hardware. Code and reconstructed scene data will be made publicly available.
Figures & tables
Figure 1: Agentic RSR overview. An agent uses real observations and known robot geometry to reconstruct a metrically aligned scene and check task-relevant interactions. A policy agent develops and tests an executable policy in simulation, then transfers it with experience memory to the real robot for feedback-driven adaptation.
Capability
CaP
Agentic R2S
CaP-X
Re 3 Sim
SimFoundry
Agentic RSR (Ours)
Real-scene reconstruction
×
✓
×
✓
✓
✓
Agentic scene revision
×
✓
×
×
×
✓
Code-policy generation
✓
×
✓
×
×
✓
Runtime agent recovery
–
×
✓
×
×
✓
Sim-to-real trials
–
–
✓
✓
✓
✓
Table 1: System-level capability comparison. Reported capabilities in real-scene reconstruction, agentic scene revision, code-policy generation, execution-time recovery, and sim-to-real trials. Symbols are defined below the table.
Figure 2: Recovering metric scale from known robot geometry. Video-estimated depth, camera parameters, and robot masks are used to construct an observed robot point cloud. Aligning this point cloud with the known URDF geometry yields metric depth and camera poses expressed in the robot-base frame.
Figure 3: Real and simulated views of a subset of task scenes.
Figure 4: Three-stage policy development. The policy progresses from privileged skill acquisition to nominal interface verification and randomized robustness, passing validated policy and experience between stages.
Method
Input
Depth MAE (m) ↓
ΔE76↓
SSIM ↑
CD (m) ↓
F1@1cm ↑
BBox (m) ↓
Re 3 Sim ( Han et al., 2025 )
Multi-view
–
20.60
0.5835
–
–
–
RL-GSBridge ( Wu et al., 2025 )
Video
–
22.87
0.5568
–
–
–
RoboSimGS ( Zhao et al., 2026 )
Multi-view
–
20.43
0.5939
–
–
–
SimFoundry ( Ranawaka et al., 2026 )
Video
0.3916
60.48
0.4495
0.1736
0.0084
0.2355
Agentic RSR (Ours)
Video
0.1057
11.04
0.6990
0.0303
0.3655
0.0869
Table 2: Scene reconstruction results across 18 tasks. – indicates that the method does not provide the outputs needed to compute the corresponding metric.
Figure 5: Representative real-robot execution frames. The initial scene is shown in (a); red crosses and green checks mark failed and successful placement outcomes in the subsequent trials.
Method
Sim SR (%) ↑
Real SR (%) ↑
Conversion (%) ↑
Code as Policies ( Liang et al., 2023 )
50.0
25.6
51.1
Instruct2Act ( Huang et al., 2023a )
38.9
27.8
71.4
Agentic RSR (Ours)
55.6
44.4
80.0
Table 3: Simulation and real-robot policy results across 18 tasks. For the Code as Policies and Instruct2Act baselines, conversion is the ratio of reported real-robot to simulation success rates, not paired task transfer. For Agentic RSR, conversion measures paired transfer among tasks that succeed in simulation.
Variant
Pass
R0.15↓
Depth MAE (m) ↓
Lab ↓
SSIM ↑
Full system
18/18
2.0
0.1057
11.04
0.6990
GPT-5.6 Sol (high)
10/18
3.9
0.2119
13.31
0.6702
w/o metric feedback
6/18
2.0
0.3746
15.12
0.6490
w/o initial calibration
3/18
6.8
0.2583
11.99
0.6597
Table 4: Real-to-Sim scene-refinement ablations across 18 scenes. R0.15 is the mean first-hit round among scenes reaching the depth threshold.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
ID
Franka task
ID
Piper task
F1
Close the drawer
P1
Close the cup lid
F2
Put corn in the bucket
P2
Put the brown block in the white-and-green storage box
F3
Put the cup in the box and close the box
P3
Put the emergency-stop button in the green storage box
F4
Stack the snacks
P4
Put the glue in the clear cup
F5
Stack two bowls
P5
Put the glue in the pen holder
F6
Store fruit in the bamboo tray
P6
Stack the brown blocks
Appendix
Table 5: Task IDs and descriptions for the nine Franka FR3 and nine Agilex Piper tasks.
Setting
Value
Sampled frames per capture
32
Rotation initialization
128 candidates; refine the top 3 and select the lowest final loss
Initial alignment
Multi-scale Sim(3) ICP
Pose refinement
Two passes over translation, rotation, and scale, followed by coarse-to-fine rotation search
Appendix
Table 6: Settings for robot-based metric-scale registration.
Robot
n
Mean Lf
Median Lf
Lf range
Mean IoU
Task-median d50 (mm)
Franka
9
0.11364
0.11221
0.10624–0.12115
0.8438
35.3
Piper
9
0.10464
0.10373
0.09997–0.11216
0.8405
21.3
Appendix
Table 7: Robot-registration results across nine captures per robot.
Method
Metric
1
2
3
4
5
6
7
8
9
Franka
Re 3 Sim
Lab
20.609
25.939
30.323
26.030
25.114
24.360
29.146
30.490
26.289
SSIM
.5612
.5334
.4875
.5280
.5229
.5028
.4980
.4708
.5151
RL-GSBridge
Lab
25.854
29.812
31.825
29.304
25.436
27.795
31.520
31.740
27.745
SSIM
.5507
.5155
.4297
.4980
.5465
.5034
.4644
.4872
.5296
RoboSimGS
Lab
20.818
24.948
28.042
24.814
23.658
23.916
26.386
28.163
24.992
Appendix
Table 8: Per-scene reconstruction metrics for Franka and Piper. Bold indicates the best comparable result among reported methods.
Variant
Task
Pass
Last round
D (m)
Lab ΔE76
SSIM
GPT-5.6 Sol
F1
No
10
0.2110
10.79
0.6883
GPT-5.6 Sol
F2
Yes
2
0.1198
10.75
0.7141
GPT-5.6 Sol
F3
No
10
0.4005
15.46
0.6718
GPT-5.6 Sol
F4
No
10
0.6543
27.77
0.5333
GPT-5.6 Sol
F5
Yes
9
0.1389
8.89
0.7757
GPT-5.6 Sol
F6
No
10
0.2516
11.46
0.7084
Appendix
Table 9: Per-scene results for three scene-refinement ablations across 18 tasks.
Adapting visuomotor policies to new manipulation tasks often requires substantial manual engineering or teleoperated data collection. Simulation can provide task-specific data at scale, but constructing the scene, designing expert behavior, and configuring data generation still require significant per-task effort. We present ARSTAG, an agentic Real2Sim2Real system that turns a single RGB image and a natural-language instruction directly into robot policy-learning data. A hierarchy of language agents constructs a task-scoped simulation scene, generates robot-feasible demonstrations, and expands the training distribution through task-consistent randomization, while a coordinator agent manages cross-stage feedback and recovery. Across seven manipulation tasks spanning grasping, placement, and stacking, the ARSTAG-generated demonstrations enable sim-to-real transfer of three visuomotor policy architectures to a dual-arm robot, with pi0.5 achieving an average real-world success rate of 74.6%. Ablations show that task-consistent randomization substantially improves robustness, and policy performance increases with generated dataset size. Project webpage: https://boweili666.github.io/ARSTAG/.
Bowei Li, Yuner Zhang, Changliu Liu
Carnegie Mellon University, Pittsburgh, PA 15213, USA
Simulated evaluation is increasingly used alongside real-world evaluation of robot policies because it is cheaper and easier to repeat; however, its value depends on how closely its outcomes track the real robot's. We test whether our reconstruction pipeline, combining metrically scaled object geometry, authored physical parameters and scene reconstruction, reduces disagreement between simulated and real robot scores relative to a default open-source recipe. We constructed two simulated versions of one bimanual robot cell: an authored reconstruction, using object geometry at estimated metric scale, projected textures, authored physics and our own scene splat; and a baseline, referred to as the default reconstruction, using the open-source recipe of a generative single-image mesh, engine-default physics and a Gaussian-splat scene. Both reconstructions use the same object photographs and scene video. Two policies ran five tasks each, giving ten task-policy pairs, which we call cells; each cell was run twenty times in each reconstruction with all other settings held fixed. Both reconstructions were scored against the same real trials, graded by a third-party evaluator. Pearson correlation between the ten simulated and real cell means is r = 0.90 for the authored reconstruction and 0.51 for the default. Mean score error is 6.97 percentage points for the authored reconstruction and 17.54 for the default, a reduction of 10.56 percentage points. These results show that improving the quality of the environment reconstruction through higher visual fidelity, authored physics and metric scale makes the simulation more faithful to the real world and narrows the sim-to-real gap. We release the harness, the per-trial scores, every reported run's configuration, and the assets and scenes of both reconstructions.
Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim process must recover scene geometries and object states, infer physical parameters, and assemble actors, objects, cameras, poses, and trajectories into a runnable physical simulation. Today this process still depends on brittle workflow glue across visual perception tools and simulators: manual tuning of visual foundation models, mesh cleanup, coordinate frame alignments, etc. We introduce \textit{Agentic Real2Sim}, a framework for generalized physical world modeling with vision-language agents that converts a real-world recording of object-robot interaction into a simulatable episodic twin, and connects the resulting twin to downstream policy fine-tuning and evaluation. We evaluate Agentic Real2Sim on rigid-object manipulation, deformable-object interaction, and humanoid motion scenes, spanning domains that are usually handled by separate Real2Sim pipelines. The framework's agentic decisions can be driven by an open-weight VLM backend at a small fraction of the cost of frontier models, while attaining a comparable conversion success rate. The framework further supports custom scene conversion, fine-tuning of a pretrained policy with data generated from converted episodes, and works effectively as a surrogate for real-world policy evaluation. The project site, including code is available at https://agentic-real2sim.github.io.
Guanxiong Chen, Qianjun Xia, Jiawei Peng +24
University of British Columbia · National University of Singapore · Johns Hopkins University +4