Coding agents powered by large language models (LLMs) have shown remarkable abilities to autonomously reason about and achieve goals in the digital world. However, bringing this success to the physical world remains challenging. On the one hand, direct generation methods (e.g., Code as Policies) often suffer from the LLMs' insufficient understanding of robots and physical environments. On the other hand, iterative trial-and-error tuning in the physical world (e.g., physical autoresearch) induces significant experimental cost and safety concerns. We introduce SimEX: Simulation-Integrated Robotics AutoResearch, an autoresearch framework that tightly integrates simulated experimentation, enabling coding agents to efficiently acquire physical capabilities for controlling real robots. SimEX operates in two stages. First, the agent conducts open-ended probe-and-optimize iterations in simulation, developing a robot toolbox with robust and generalizable capabilities. Second, the agent adapts the toolbox and the simulator together through only a few physical trials: each trial corrects the simulator, and the corrected simulator is used to diagnose failures and screen candidate repairs. We evaluate SimEX extensively in sim-to-sim settings and on physical robots. On challenging real-world manipulation tasks including towel folding, barcode scanning, and plate manipulation, SimEX enables coding agents to efficiently acquire robot skills without any demonstration and with only 10 minutes of real-robot interaction. These results suggest that simulation can be a critical component in achieving physical intelligence, not only as a source of training data that must closely replicate the real world, but also as a roughly correct laboratory where a coding agent develops the knowledge and procedures needed to act on the robot. More details and robot videos at https://robo-simex.github.io/
Figures & tables
Figure 1: SimEX brings coding agents to physical robots through simulation. (1) The agent develops diverse robot capabilities through open-ended autoresearch in simulation. (2) The agent then efficiently adapts to the real robot within a few trials, turning each real rollout into many simulated tests of candidate repairs. Bottom: physical-robot rollouts using the optimized toolboxes.
Figure 2: Overview of SimEX. Preparation reconstructs the physical workspace in simulation and initializes the toolbox. Stage 1 runs independent probe–focus–optimize cycles, retaining improvements on focused tasks and capability-bank replay. Stage 2 executes one robot rollout, aligns the simulator through fixed-policy replay, creates candidate repairs, and returns simulated screening evidence to the coding agent as supportive evidence for repair selection.
Figure 3: Aggregate performance across the three evaluation task families.
Figure 4
Method
Plate
Towel
Barcode
Direct CaP
0/10
0/10
0/10
ENPIRE ∗
2/10
1/10
0/10
ZS Sim2Real
0/10
2/10
0/10
ASPIRE ∗
2/10
0/10
0/10
SimEX (ours)
10/10
8/10
8/10
Table 1: Physical-robot successful evaluation trials out of 10.
Figure 6: Cross-agent comparison. Our main conclusion holds with different coding agents.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Phase
Hyperparameter
Value
Stage 1 & 2
Coding-agent orchestrator
Fable 5.1
Stage 1 & 2
Policy-writing model
Fable 5.1
Stage 1
Reasoning effort
max
Stage 2
Reasoning effort
medium
Stage 1
Independent population members
4
Stage 1
Optimization iterations per member
20
Appendix
Table 2: Default hyperparameters used in the main experiments. Population size, per-process pretraining budget, deployment-trial budget, and evaluation protocol are matched across methods whenever the corresponding phase applies.
Stage 1 (per member)
Stage 2 (per run)
Orchestrator calls
840
124
Agent thinking time
5.4 h
31 min
Output tokens (reasoning)
1.98M (1.16M)
161k (64k)
Cache-read input tokens
155M
19.5M
Appendix
Table 3: Coding agent orchestration cost.
Figure 7: Physical setups for the three evaluation task families: plate to tote, towel folding, and barcode scanning. All use the same dual-arm YAM platform and workspace enclosure, while task-specific objects and fixtures produce distinct manipulation challenges.
Figure 8: Initial states of the seven plate to tote tasks. (Isaac Sim)
Figure 9: Initial states of the seven towel folding tasks. (mjlab)
Figure 10: Initial states of the seven barcode scanning tasks. (Isaac Sim)
Scene
Task
Direct CaP
ENPIRE ∗
Zero-Shot Sim2Real
ASPIRE ∗
SimEX (no discovery)
SimEX (direct repair)
SimEX (ours)
Plate to tote
Same-side pair
1/25
4/25
1/25
0/25
16/25
21/25
21/25
Diagonal pair
3/25
8/25
0/25
0/25
12/25
14/25
24/25
Far outboard plate
17/25
14/25
12/25
17/25
15/25
19/25
22/25
Far plate to near tote
9/25
18/25
10/25
19/25
20/25
25/25
25/25
Any three
0/25
6/25
0/25
0/25
23/25
16/25
23/25
Shifted single
11/25
5/25
25/25
16/25
6/25
23/25
25/25
Appendix
Table 4: Per-task controlled sim-to-sim results, reported as successful evaluation trials out of 25.
Task
SimEX (ours)
No population (Stage 1)
No capability bank (Stage 1)
No repair alternatives (Stage 2)
No hypothesis generation (Stage 2)
Wall-side bottle
25/25
17/25
15/25
24/25
23/25
Across-tote bottle
24/25
20/25
14/25
25/25
20/25
Far upright bottle
24/25
25/25
25/25
25/25
20/25
Hidden-barcode block
24/25
25/25
24/25
25/25
25/25
Wall-side prism
19/25
15/25
21/25
5/25
13/25
Two bottles
19/25
12/25
15/25
18/25
16/25
Appendix
Table 5: Per-task barcode scanning ablations, reported as successful evaluation trials out of 25.
Figure 11: Across all coding agents, SimEX consistently outperforms the strongest baseline methods (marked with darker diamond) by a large margin.
Task
Direct CaP
ENPIRE ∗
Zero-Shot Sim2Real
ASPIRE ∗
SimEX (no discovery)
SimEX (direct repair)
SimEX (ours)
Wall-side bottle
0/25
0/25
0/25
0/25
2/25
17/25
25/25
Across-tote bottle
0/25
0/25
0/25
1/25
7/25
17/25
24/25
Far upright bottle
0/25
0/25
0/25
15/25
17/25
16/25
24/25
Hidden-barcode block
1/25
0/25
0/25
10/25
8/25
25/25
24/25
Wall-side prism
0/25
0/25
0/25
6/25
6/25
3/25
19/25
Two bottles
0/25
0/25
2/25
0/25
2/25
20/25
19/25
Appendix
Table 6: Per-task barcode scanning results for the all-Fable 5.1 configuration used in the main experiments, reported as successful evaluation trials out of 25.
Task
Direct CaP
ENPIRE ∗
Zero-Shot Sim2Real
ASPIRE ∗
SimEX (no discovery)
SimEX (direct repair)
SimEX (ours)
Wall-side bottle
0/25
4/25
0/25
0/25
0/25
16/25
25/25
Across-tote bottle
0/25
1/25
0/25
0/25
3/25
20/25
25/25
Far upright bottle
0/25
8/25
0/25
0/25
7/25
20/25
22/25
Hidden-barcode block
0/25
2/25
0/25
0/25
17/25
24/25
23/25
Wall-side prism
0/25
0/25
0/25
0/25
4/25
2/25
10/25
Two bottles
0/25
0/25
3/25
0/25
1/25
21/25
23/25
Appendix
Table 7: Per-task barcode scanning results for the end-to-end Opus 5 configuration, reported as successful evaluation trials out of 25.
Task
Direct CaP
ENPIRE ∗
Zero-Shot Sim2Real
ASPIRE ∗
SimEX (no discovery)
SimEX (direct repair)
SimEX (ours)
Wall-side bottle
0/25
0/25
0/25
0/25
2/25
19/25
21/25
Across-tote bottle
0/25
0/25
0/25
9/25
1/25
13/25
21/25
Far upright bottle
0/25
0/25
0/25
1/25
0/25
10/25
17/25
Hidden-barcode block
0/25
0/25
0/25
5/25
15/25
25/25
24/25
Wall-side prism
0/25
0/25
0/25
5/25
4/25
0/25
5/25
Two bottles
0/25
0/25
7/25
0/25
0/25
25/25
24/25
Appendix
Table 8: Per-task barcode scanning results for the end-to-end GPT-5.5 configuration, reported as successful evaluation trials out of 25.
Task
Direct CaP
ENPIRE ∗
Zero-Shot Sim2Real
ASPIRE ∗
SimEX (no discovery)
SimEX (direct repair)
SimEX (ours)
Wall-side bottle
0/25
0/25
0/25
0/25
4/25
17/25
25/25
Across-tote bottle
0/25
0/25
0/25
0/25
0/25
17/25
24/25
Far upright bottle
0/25
0/25
0/25
0/25
6/25
20/25
25/25
Hidden-barcode block
0/25
0/25
0/25
10/25
9/25
25/25
25/25
Wall-side prism
0/25
0/25
0/25
0/25
0/25
0/25
8/25
Two bottles
0/25
0/25
0/25
0/25
0/25
17/25
20/25
Appendix
Table 9: Per-task barcode scanning results for the end-to-end GPT-6 Astra configuration, reported as successful evaluation trials out of 25.
Achieving dexterous robotic manipulation in the real world heavily relies on human supervision and algorithm engineering, which becomes a central bottleneck in the pursuit of general physical intelligence. Although emerging coding agents can generate code to automate algorithm search, their successes remain largely confined in digital environments. We conjecture that the missing abstraction to automate robotics research is a repeatable feedback loop for real-world policy improvement: reset the scene, execute a policy, verify the outcome, and refine the next iteration. To bridge this gap, we introduce ENPIRE, a harness framework for coding agents that instantiates this physical feedback routine with four core modules: an Environment module (EN) for automatic reset and verification, a Policy Improvement module (PI) that launches policy refinement, a Rollout module (R) to evaluate policies with one or multiple physical robots operating in parallel, and an Evolution module (E) in which coding agents analyze logs, consult literature, improve training infrastructure and algorithm code to address failure modes. This closed-loop system transforms real-world manipulation learning into a controllable optimization procedure, minimizing human effort while allowing fair ablations across training recipe and agent variants. Powered by ENPIRE, frontier coding agents can autonomously train a policy to achieve a 99% success rate on challenging, dexterous manipulation tasks, such as organizing a pin box, fastening a zip tie, and tool use, a process that further accelerates when we dispatch an agent team on a robot fleet. Our results suggest a practical and scalable path toward deploying coding agents to autonomously advancing robotics in the physical world.
Developing navigation policies requires simulation scenarios that support repeatable training and evaluation. Despite the availability of numerous simulators, constructing diverse navigation scenarios often involves writing simulator-specific code, using complex graphical interfaces, performing substantial manual configuration, or relying on resource-intensive computing platforms, making scenarios difficult to reproduce and reuse across experiments. To this end, we develop the Intelligent Robot Simulator (IR-SIM), a lightweight declarative simulator implemented as a Python library to support navigation learning and benchmarking through rapid construction of reusable and diverse scenarios and efficient execution on accessible platforms. IR-SIM represents scenarios as human-readable YAML configurations, in which necessary objects, behaviors, sensors, maps, and environment parameters can be flexibly specified and composed. This declarative representation makes IR-SIM friendly to large language model (LLM)-powered agents, enabling them to compose and modify executable navigation scenario files through the provided agent skills, instead of writing a large amount of code that is difficult to reuse. Despite its lightweight implementation, IR-SIM provides the essential components for navigation simulation, while retaining interfaces to high-fidelity simulators for downstream validation. Experiments demonstrate that IR-SIM runs up to several hundred times faster on the evaluated CPU platform, while agent skills reduce mean LLM-based scenario construction time by more than 50% across both evaluated models. The experiments further show ability to support reproducible benchmarking and RL-based navigation policy learning.
Ruihua Han, Shuai Wang, Chengyang Li +8
The University of Hong Kong · Shenzhen Institutes of Advanced Technology · Southern University of Science and Technology +2
Real-world videos provide rich demonstrations of manipulation, but turning them into reusable robot skills requires visually aligned environments, executable physical interactions, and mechanisms for learning from experience. We introduce Real2Gym, an agentic Real2Sim2Real framework that turns human and robot demonstrations into interactive simulation gyms and brings skills acquired in simulation to physical robots. The Real2Sim module reconstructs editable scenes, aligns objects and cameras with the input, validates demonstrated or retargeted actions through native physics execution, and generates task-conditioned variations with action-feasibility checks. Within these environments, the agent generates executable code for manipulation stages, observes their outcomes, and distills successful attempts and failures into reusable task procedures, object-relative motions, and recovery strategies. Through a shared perception-and-control interface, these skills guide subsequent execution in simulation and on real robots, with motions adapted to current observations and no updates to the underlying model weights. Extensive evaluations demonstrate that Real2Gym enables high-fidelity simulation environment reconstruction, outperforming GPT-6 Astra Direct Mode by 16.7% in success rate with approximately 74.9% fewer policy-execution tokens across these environments, while exceeding it by 33.3% in physical robot execution success rate across four tasks on a real Franka robot.
Kerui Ren, Yingxiang Xu, Kaiwen Song +4
Shanghai Artificial Intelligence Laboratory · Shanghai Jiao Tong University · Zhejiang University +3