Organizations: Eastern Institute of Technology, Ningbo, China · National University of Singapore, Singapore · Department of Aeronautical and Aviation Engineering, The Hong Kong Polytechnic University, Hong Kong · University of Science and Technology of China, Hefei, China
Simulation-to-robot transfer can fail when velocity commands produce motion and feedback that differ from those modeled during policy training. We present execution-interface dynamics adaptation (EIDA), which fits these responses from target-platform execution data without reconstructing actuator dynamics. A model of body-frame pose increments updates simulator geometry, while a separate model predicts the velocity feedback observed by the policy; a short history of velocity feedback is included in the policy input. The fitted models are used within a lightweight GPU-parallel simulator. On the full Jackal and Go2 validation sets, the fitted models reduced position and yaw prediction errors relative to the simulator's predefined motion model. Across 100 benchmark navigation environments evaluated in a separate physics-based simulator, EIDA achieved the highest success rate and navigation score among the compared learned policies, both with and without global guidance. Feedback ablations further supported the need to match policy-facing velocity estimates. On a physical Unitree Go2, EIDA reached the goal without collision in all 20 static-scene trials, compared with 4 of 20 for the baseline. These results show that execution-interface adaptation can improve navigation transfer without detailed actuator simulation.
Figures & tables
Fig. 1: Sim-to-real execution mismatch can cause navigation failures. EIDA incorporates models learned from real executions into lightweight simulation to improve navigation policy transfer to physical robots.
Fig. 2: Execution-interface dynamics adaptation (EIDA). Target-system execution data identify independent pose-increment and velocity-feedback models. The execution-fitted interface supports policy training in simulation; at deployment, live sensing and the existing controller replace the fitted responses.
Fig. 3: Pose-prediction accuracy on the Jackal and Go2 validation sets. Curves show mean endpoint RMSE over the prediction horizon; error bars indicate one standard deviation across three pose-model seeds.
Fig. 4: Training and evaluation settings: (a) Simulation training map, (b) Jackal in BARN, and (c) physical Go2.
Method
SR (%) ↑
CR (%) ↓
TR (%) ↓
Metric ↑
Global
DWA
46.00
37.00
17.00
0.2130
Yes
E-Band
57.70
9.70
32.60
0.2880
Yes
Nominal
71.22 ± 5.68
26.89
1.89
0.3547
No
Nominal
80.78 ± 1.35
18.44
0.78
0.3971
Yes
RWM-U
69.56 ± 4.40
25.56
4.89
0.3469
No
RWM-U
73.78 ± 5.67
25.56
0.67
0.3680
Yes
TABLE I: Simulation-based performance comparison of different models using the BARN dataset.
Fig. 5: Trajectories of the robot controlled by (a) Nominal, (b) RWM-U, and (c) EIDA in one of the 100 BARN testing environments. “S” represents the starting point, and “G” represents the target point. The corresponding videos are provided in the supplementary material.
Model
Hist.
Feedback
Global path: No
Global path: Yes
SR (%)
Metric
SR (%)
Metric
Nominal
No
Command
71.67
0.3555
81.00
0.3977
Nominal
Yes
Command
70.67
0.3525
83.33
0.4082
Nominal
No
ARX
70.67
0.3505
83.33
0.4109
Nominal
Yes
ARX
72.33
0.3611
76.00
0.3726
Fitted
No
Command
73.33
0.3622
82.00
0.3954
TABLE II: Ablation of motion modeling, velocity history, and feedback source on BARN.
Fig. 6: Physical Go2 navigation in two static scenes. Panels (a)–(d) compare the Nominal baseline with EIDA; a dynamic scenario in which a pedestrian suddenly obstructs the path is shown separately in Fig. 7 .
Fig. 7: Go2 navigation with EIDA in Scene 2 when a pedestrian suddenly obstructs the path. Overlaid robot and pedestrian poses illustrate the motion sequence toward the marked goal. This demonstration is separate from the static-scene trials in Table III .
Policy
Collision ↓
Goal ↑
Path (m) ↓
Time (s) ↓
Scene 1: Curved passage
Nominal
9/10
1/10
6.81
12.93
EIDA
0/10
10/10
6.32 ± 0.18
10.31 ± 0.75
Scene 2: Separated obstacles
Nominal
7/10
3/10
6.71 ± 0.37
12.04 ± 2.29
EIDA
0/10
10/10
5.91 ± 0.30
8.98 ± 1.04
TABLE III: Real-world Go2 navigation performance in two scenes.
Sim-to-real transfer has made substantial progress, but can still produce controllers that remain stable and functional on hardware while suffering from degraded tracking accuracy due to residual dynamics mismatch. Correcting these errors typically requires identifying the underlying system dynamics, adapting the control policy, or returning to simulation for additional training and finetuning, all of which can require substantial data and computation. We propose OSRAM (Online Sim-to-Real Adaptation via Closed-Loop System Modeling), a framework that instead adapts the reference commands provided to an existing controller. OSRAM treats the deployed robot and its policy as a unified closed-loop dynamical system and learns its task-level command-response behavior directly from tracking observations. A closed-loop dynamics model is meta-trained across randomized dynamics in simulation and rapidly finetuned after deployment using limited real-world interaction. The adapted model is then used to optimize future reference commands while leaving the underlying control policy unchanged. We evaluate OSRAM on bipedal velocity tracking and loco-manipulation in simulation and on hardware. Results show that closed-loop modeling improves prediction and tracking accuracy under unseen dynamics, while online reference adaptation reduces residual sim-to-real tracking errors across different control objectives and hardware configurations. These results demonstrate that adapting the behavior of the robot-policy closed loop provides a practical alternative to finetuning the policy or identifying the full physical dynamics for sim-to-real transfer. More information can be found at http://generalroboticslab.com/OSRAM.
In recent years, reinforcement learning (RL) has shown remarkable success in robotics when a fast and accurate simulator is available for a given task. When using RL and simulation, more simulator realism is generally beneficial but becomes harder to obtain as robots are deployed in increasingly complex and widescale domains. In such settings, simulators will likely fail to model all relevant details of a given target task and this observation motivates the study of sim2real with simulators that leave out key task details. In this paper, we formalize and study the abstract sim2real problem: given an abstract simulator that models a target task at a coarse level of abstraction, how can we train a policy with RL in the abstract simulator and successfully transfer it to the real-world? Our first contribution is to formalize this problem using the language of state abstraction from the RL literature. This framing shows that an abstract simulator can be grounded to match the target task if the grounded abstract dynamics take the history of states into account. Based on the formalism, we then introduce a method that uses real-world task data to correct the dynamics of the abstract simulator. We then show that this method enables successful policy transfer both in sim2sim and sim2real evaluation.
Yunfu Deng, Yuhao Li, Josiah P. Hanna
Department of Computer Sciences, University of Wisconsin–Madison, Madison, WI 53706 USA · Manning College of Information and Computer Sciences, University of Massachusetts Amherst, Amherst, MA 01003 USA
Adapting robot manipulation policies to new tasks and environments remains highly data-intensive, while the data needed for further improvement depends on the policy's current capabilities and failure modes. We introduce EmbodiRSI, an agentic system for recursive self-improvement (RSI) in a real-to-sim-to-real setting, where task-specific simulations are constructed from target deployment scenarios and used as low-cost environments for iterative policy improvement before transfer back to the physical world. EmbodiRSI uses policy execution feedback to guide subsequent experience acquisition and policy updates. Two complementary mechanisms close this loop: Collaborative Error Correction generates agent-assisted corrective trajectories from policy-reached states, while Adaptive Data Collection directs expert demonstration generation toward the current policy's weaknesses. The task-specific simulation serves as a reusable workspace for policy warm-up, repeatable evaluation, failure diagnosis, and targeted data generation across successive RSI rounds. Across three tabletop environments and 14 subtasks, EmbodiRSI increases scene-balanced autonomous simulation success from 50.4% to 83.5% over two RSI updates. With 400 adaptive simulated trajectories and only ten real-world refinement trajectories per subtask, EmbodiRSI achieves 83.1% scene-balanced autonomous real-world success, compared with 75.0% for adaptation using 200 real-world demonstrations per subtask. These results demonstrate that feedback-driven recursive improvement in deployment-specific simulations can enable data-efficient adaptation of embodied policies to physical environments.