cs.ROOct 8, 2026

Leveraging Human-In-The-Loop Demonstrations in Reinforcement Learning for Digital Twin-Driven Robot Flexibility

Authors: Yuzhu Sun, Mien Van, Nguyen Minh Nhat, Stephen McIlvanna, Sean McLoone

Organizations: School of Electronics, Electrical Engineering and Computer Science, Queen’s University Belfast, Northern Ireland, UK

Abstract

Growing automation makes collaborative robots work in more variable environments, increasing the need for adaptation. We propose a human-in-the-loop online training framework combining a digital twin (DT), reinforcement learning (RL), and human demonstrations. Unlike DTs used mainly to generate synthetic data before task execution, our DT is synchronized with the physical system in real time through camera feeds, allowing the virtual robot to update its observations and policy from real-world feedback. A dual actor framework integrates imitation learning (IL) without adding a direct imitation loss to the RL actor, so demonstrations can guide adaptation instead of manual reprogramming. The proposed framework is demonstrated on the Ufactory Xarm5 collaborative robot, where the robot's end-effector aims to reach the target position while avoiding obstacles. The experiments show that the framework can resume training after a change in the physical workspace and that, with a fixed set of non-optimal demonstrations, the dual actor framework achieves a much higher final success rate than two methods that add an imitation loss to the actor. The same pattern holds with real human demonstrations collected in virtual reality (VR): with demonstrations that never reach the goal, the dual actor framework reached 83-100% mean deterministic evaluation success, against 0-17% for the two imitation-loss methods.

Figures & tables

Explore similar work

CardsList
  1. One Demonstration Is Enough for Real-World Robotic Reinforcement Learning

    Jul 2, 2026Yuwan Liu, Hongze Yu, Song Liu +5Reinforcement LearningRobot Skill Learning

  2. Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation

    Sep 24, 2026Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht +1Contact-Rich Robotic ManipulationHuman-in-the-Loop AI

  3. RoHIL: Robust Human-in-the-Loop Robotic Reinforcement Learning Against Illumination Variations

    May 19, 2026Shuoqin Zhang, Yixin Xiong, Xiru Gao +4Distribution Shift RobustnessRobot Skill Learning