A simulation of a real robot workspace must preserve task-relevant interactions, while policies developed in it must operate on observations available to the real robot. Yet scene reconstruction and policy development are often treated separately. We present Agentic Real-to-Sim-to-Real (Agentic RSR), a framework that links scene reconstruction, policy development, and real-robot execution through the same manipulation task. Given a workspace video, a task description, and a known robot model, an agent recovers metric scale, iteratively refines the scene using visual feedback, and checks task-relevant interactions in MuJoCo. A coding agent then develops an executable policy, progressing from privileged object poses to visual observations and randomized simulation. The policy can interleave multiple observations and actions within one invocation, while the agent uses execution feedback to continue, retry, or revise its approach. A shared task-level interface carries the policy and accumulated experience to the real robot, where fresh observations and safety checks guide execution. Across 18 reconstructed scenes involving two robots, the mean four-view Depth MAE against reference depth estimates is 0.1057 m, the mean Lab ΔE76 is 11.04, and the mean grayscale SSIM is 0.6990. In real-robot experiments, the aggregate task success rate reaches 80% of the simulation task success rate, indicating substantial retention of simulated performance on hardware. Code and reconstructed scene data will be made publicly available.
Figures & tables
Figure 1: Agentic RSR overview. An agent uses real observations and known robot geometry to reconstruct a metrically aligned scene and check task-relevant interactions. A policy agent develops and tests an executable policy in simulation, then transfers it with experience memory to the real robot for feedback-driven adaptation.
Capability
CaP
Agentic R2S
CaP-X
Re 3 Sim
SimFoundry
Agentic RSR (Ours)
Real-scene reconstruction
×
✓
×
✓
✓
✓
Agentic scene revision
×
✓
×
×
×
✓
Code-policy generation
✓
×
✓
×
×
✓
Runtime agent recovery
–
×
✓
×
×
✓
Sim-to-real trials
–
–
✓
✓
✓
✓
Table 1: System-level capability comparison. Reported capabilities in real-scene reconstruction, agentic scene revision, code-policy generation, execution-time recovery, and sim-to-real trials. Symbols are defined below the table.
Figure 2: Recovering metric scale from known robot geometry. Video-estimated depth, camera parameters, and robot masks are used to construct an observed robot point cloud. Aligning this point cloud with the known URDF geometry yields metric depth and camera poses expressed in the robot-base frame.
Figure 3: Real and simulated views of a subset of task scenes.
Figure 4: Three-stage policy development. The policy progresses from privileged skill acquisition to nominal interface verification and randomized robustness, passing validated policy and experience between stages.
Method
Input
Depth MAE (m) ↓
ΔE76↓
SSIM ↑
CD (m) ↓
F1@1cm ↑
BBox (m) ↓
Re 3 Sim ( Han et al., 2025 )
Multi-view
–
20.60
0.5835
–
–
–
RL-GSBridge ( Wu et al., 2025 )
Video
–
22.87
0.5568
–
–
–
RoboSimGS ( Zhao et al., 2026 )
Multi-view
–
20.43
0.5939
–
–
–
SimFoundry ( Ranawaka et al., 2026 )
Video
0.3916
60.48
0.4495
0.1736
0.0084
0.2355
Agentic RSR (Ours)
Video
0.1057
11.04
0.6990
0.0303
0.3655
0.0869
Table 2: Scene reconstruction results across 18 tasks. – indicates that the method does not provide the outputs needed to compute the corresponding metric.
Figure 5: Representative real-robot execution frames. The initial scene is shown in (a); red crosses and green checks mark failed and successful placement outcomes in the subsequent trials.
Method
Sim SR (%) ↑
Real SR (%) ↑
Conversion (%) ↑
Code as Policies ( Liang et al., 2023 )
50.0
25.6
51.1
Instruct2Act ( Huang et al., 2023a )
38.9
27.8
71.4
Agentic RSR (Ours)
55.6
44.4
80.0
Table 3: Simulation and real-robot policy results across 18 tasks. For the Code as Policies and Instruct2Act baselines, conversion is the ratio of reported real-robot to simulation success rates, not paired task transfer. For Agentic RSR, conversion measures paired transfer among tasks that succeed in simulation.
Variant
Pass
R0.15↓
Depth MAE (m) ↓
Lab ↓
SSIM ↑
Full system
18/18
2.0
0.1057
11.04
0.6990
GPT-5.6 Sol (high)
10/18
3.9
0.2119
13.31
0.6702
w/o metric feedback
6/18
2.0
0.3746
15.12
0.6490
w/o initial calibration
3/18
6.8
0.2583
11.99
0.6597
Table 4: Real-to-Sim scene-refinement ablations across 18 scenes. R0.15 is the mean first-hit round among scenes reaching the depth threshold.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
ID
Franka task
ID
Piper task
F1
Close the drawer
P1
Close the cup lid
F2
Put corn in the bucket
P2
Put the brown block in the white-and-green storage box
F3
Put the cup in the box and close the box
P3
Put the emergency-stop button in the green storage box
F4
Stack the snacks
P4
Put the glue in the clear cup
F5
Stack two bowls
P5
Put the glue in the pen holder
F6
Store fruit in the bamboo tray
P6
Stack the brown blocks
Appendix
Table 5: Task IDs and descriptions for the nine Franka FR3 and nine Agilex Piper tasks.
Setting
Value
Sampled frames per capture
32
Rotation initialization
128 candidates; refine the top 3 and select the lowest final loss
Initial alignment
Multi-scale Sim(3) ICP
Pose refinement
Two passes over translation, rotation, and scale, followed by coarse-to-fine rotation search
Appendix
Table 6: Settings for robot-based metric-scale registration.
Robot
n
Mean Lf
Median Lf
Lf range
Mean IoU
Task-median d50 (mm)
Franka
9
0.11364
0.11221
0.10624–0.12115
0.8438
35.3
Piper
9
0.10464
0.10373
0.09997–0.11216
0.8405
21.3
Appendix
Table 7: Robot-registration results across nine captures per robot.
Method
Metric
1
2
3
4
5
6
7
8
9
Franka
Re 3 Sim
Lab
20.609
25.939
30.323
26.030
25.114
24.360
29.146
30.490
26.289
SSIM
.5612
.5334
.4875
.5280
.5229
.5028
.4980
.4708
.5151
RL-GSBridge
Lab
25.854
29.812
31.825
29.304
25.436
27.795
31.520
31.740
27.745
SSIM
.5507
.5155
.4297
.4980
.5465
.5034
.4644
.4872
.5296
RoboSimGS
Lab
20.818
24.948
28.042
24.814
23.658
23.916
26.386
28.163
24.992
Appendix
Table 8: Per-scene reconstruction metrics for Franka and Piper. Bold indicates the best comparable result among reported methods.
Variant
Task
Pass
Last round
D (m)
Lab ΔE76
SSIM
GPT-5.6 Sol
F1
No
10
0.2110
10.79
0.6883
GPT-5.6 Sol
F2
Yes
2
0.1198
10.75
0.7141
GPT-5.6 Sol
F3
No
10
0.4005
15.46
0.6718
GPT-5.6 Sol
F4
No
10
0.6543
27.77
0.5333
GPT-5.6 Sol
F5
Yes
9
0.1389
8.89
0.7757
GPT-5.6 Sol
F6
No
10
0.2516
11.46
0.7084
Appendix
Table 9: Per-scene results for three scene-refinement ablations across 18 tasks.