An Real-Sim-Real (RSR) Loop Framework for Generalizable Robotic Policy Transfer
Authors: Yuxuan Xu, Shiyu Wang, Jinhao Huang, Wenhao Zhao, Yufei Jia, Zike Yan, Weibin Gu, Lu Shi, +1 more
Organizations: Institute for AI Industry Research (AIR), Tsinghua University · Beijing Jiaotong University · The Hong Kong University of Science and Technology (Guangzhou)
The sim-to-real gap remains a critical challenge in robotics, hindering the deployment of algorithms trained in simulation to real-world systems. We propose a flexible Real-to-Sim-to-Real (RSR) framework whose central contribution is an information-theoretic cost function that explicitly accounts for sim-to-real discrepancies. This cost balances two objectives, completing the task and steering the policy to collect real-world samples that are maximally informative for improving transfer. It can be integrated seamlessly into existing reinforcement learning algorithms (e.g., PPO, SAC) and ensures a balanced exploration of critical regions in the real domain. The framework treats differentiable simulation as optional: when a differentiable simulator is available, the collected informative data can also be used to tune simulator parameters. We implement the framework with the MuJoCo MJX platform and demonstrate its generality by evaluating on both manipulation tasks with a 6-DOF robotic arm and locomotion tasks on a legged robot. Empirical results show that our RSR loop yields more efficient data acquisition and substantially improves task performance in real-world that achieves a smoother sim-to-real transfer.
Figures & tables
Fig. 1: Overview of the proposed RSR (Real–Sim–Real) loop framework. The red outer loop (arrows labeled “R”) is the main Real–Sim–Real cycle: a simulation-trained policy is deployed on the real robot, the collected data are used to update the InfoGap cost (and, optionally, the simulator), and a new policy is then trained in simulation. The blue inner loop is the policy-training loop, which uses the current simulator and the adaptive InfoGap loss. The green inner loop is the optional simulator-parameter-tuning loop (indexed by i). Together, these loops iteratively improve both the policy and, when applicable, the simulator.
θi,k−1←θi−1,k−1−α∇θL(θi−1,k−1)
Algorithm 1 The RSR Loop Framework
Fig. 2: The process of bridging the sim-to-real gap in robot training. When the discrepancy between the simulation (blue domain) and real robot (orange domain) is large, the policy prioritizes collecting informative data (marked as crosses) from the real domain to better characterize its properties other than the task trajectory (black dashed line).
Stage
KL Divergence ( x )
KL Divergence ( y )
PPO+DR
16.3509
36.3168
1st_RSR
1.6903
3.4206
2nd_RSR
5.0462
2.3946
3rd_RSR
0.9739
0.8384
TABLE I: Distribution Deviation Results in the Block-Pushing Task
Fig. 3: Real-world block-pushing trajectories across different iterations.
Fig. 4: 1- σ bounds of real trajectories for the yaw angle in the T-shaped block pushing trials for different iterations, where the shaded area represents the bounds.
Stage
KL Divergence
vx (m/s)
vy (m/s)
ωz (rad/s)
PPO+DR
48.2941
5.1624
56.7724
1st_RSR
38.3659
3.5171
47.5421
2nd_RSR
9.2554
2.6966
19.5236
3rd_RSR
4.3988
1.9778
3.9715
TABLE II: Distribution Deviation of the Circular-Tracking Task
Fig. 5: Real-world legged trajectories across different iterations.
Stage
vx (m/s)
vy (m/s)
ωz (rad/s)
PPO+DR
0.194
0.077
0.334
1st_RSR
0.147
0.050
0.295
2nd_RSR
0.067
0.068
0.086
3rd_RSR
0.051
0.045
0.052
TABLE III: Velocity Tracking RMSE of the Quadrupedal Robot
Payload
Stage
vx (m/s)
vy (m/s)
ωz (rad/s)
None
SAC+DR
44.8173
4.9362
51.2847
1st RSR
30.9648
4.0837
35.6724
2nd RSR
17.3826
2.9145
19.4861
3rd RSR
8.7519
1.7286
8.9375
Centered
SAC+DR
52.4618
6.2743
63.9052
1st RSR
37.5831
4.8619
44.7168
TABLE IV: Distribution Deviation across RSR Iterations under Different Payload Conditions
Fig. 6: Velocity-tracking RMSE under different payload conditions across successive RSR and SAC+Real Replay iterations. Both methods use the same real-world transition sets at each adaptation round.
Fig. 7: Trajectories of real-robot deployment with and without visual loss.
In recent years, reinforcement learning (RL) has shown remarkable success in robotics when a fast and accurate simulator is available for a given task. When using RL and simulation, more simulator realism is generally beneficial but becomes harder to obtain as robots are deployed in increasingly complex and widescale domains. In such settings, simulators will likely fail to model all relevant details of a given target task and this observation motivates the study of sim2real with simulators that leave out key task details. In this paper, we formalize and study the abstract sim2real problem: given an abstract simulator that models a target task at a coarse level of abstraction, how can we train a policy with RL in the abstract simulator and successfully transfer it to the real-world? Our first contribution is to formalize this problem using the language of state abstraction from the RL literature. This framing shows that an abstract simulator can be grounded to match the target task if the grounded abstract dynamics take the history of states into account. Based on the formalism, we then introduce a method that uses real-world task data to correct the dynamics of the abstract simulator. We then show that this method enables successful policy transfer both in sim2sim and sim2real evaluation.
Yunfu Deng, Yuhao Li, Josiah P. Hanna
Department of Computer Sciences, University of Wisconsin–Madison, Madison, WI 53706 USA · Manning College of Information and Computer Sciences, University of Massachusetts Amherst, Amherst, MA 01003 USA
To mitigate the sample complexity of real-world reinforcement learning (RL), a common practice is to first train a policy in a simulator, where samples are cheap, and then deploy the learned policy in the real world with the hope that it generalizes effectively. Such direct sim-to-real transfer is not guaranteed to succeed: simulator-trained policies can be suboptimal in the real world due to sim-to-real mismatch. Correcting this mismatch requires collecting data from the real system, but in many applications, such as robotics and healthcare, this data-collection process is itself subject to safety constraints. This gives rise to the problem of safe sim-to-real transfer: how can an agent exploit an imperfect simulator while ensuring safe real-world data collection and learning a near-optimal feasible policy for the target system? We address this problem by formulating safe sim-to-real transfer within the framework of reward-free safe RL. We design a computationally efficient algorithm that exploits simulator information to provably reduce real-world interaction while ensuring safe exploration and enabling the computation of a near-optimal feasible policy for any potential reward function. Our real-world sample complexity bound characterizes the benefit of using the simulator in terms of the sim-to-real mismatch.
Robots trained on real world data tend to be imprecise, slow, and brittle to perturbations. Improving these policies with reinforcement learning (RL) is an appealing alternative, but this process often requires expensive training in the real world. Performing policy improvement in simulation instead provides a far cheaper alternative, but unconstrained RL in simulation can exploit contact and dynamics mismatches, resulting in unsafe behaviors that do not transfer to hardware. Common forms of regularization can furthermore limit improvement by overconstraining to an imperfect behavior prior. In this work, we propose Support-Constrained Off-Domain REinforcement (SCORE), a real-to-sim-to-real framework that constrains RL in simulation to the support of a generative policy pretrained on real data. We instantiate this constraint through flow steering, restricting SCORE to actions the base policy can already produce, which ensures transferable behaviors while maximizing policy improvement. Improving a policy with SCORE requires minimal effort: it learns from sparse rewards, avoids distillation, and leaves the base policy untouched. Across eight real-world dexterous multi-fingered robotic manipulation tasks, SCORE improves average success rate from 37.8% to 89.9%, compared to 59.5% for the best baseline, and reaches success in 36.8% fewer steps than the base policy. Ultimately, through extensive experiments and ablations, we show that simulation can substantially improve real-world manipulation policies when policy optimization is appropriately constrained, introducing a new paradigm for real-to-sim-to-real policy improvement. Videos and code are available at https://weirdlabuw.github.io/score/.