cs.ROSep 24, 2026

Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation

Authors: Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv

Organizations: Siemens AG, Research and Predevelopment, Germany. · Technical University of Munich, Germany.

Abstract

Imitation learning enables robots to acquire manipulation skills from demonstrations, but the resulting policies can fail outside the training data, while collecting more demonstrations requires substantial human effort. Human-in-the-loop reinforcement learning uses corrective feedback during online training, but typically learns the complete task policy rather than refining a pretrained imitation policy. We introduce Res-HIL, a human-in-the-loop residual reinforcement learning framework that learns corrective actions on top of a frozen imitation policy. Each human intervention provides two complementary learning signals: direct supervision of the residual policy and reward shaping of preceding autonomous behavior. Res-HIL combines these signals with zero initialization of the residual policy to stabilize and accelerate online learning. We evaluate Res-HIL on five contact-rich manipulation tasks spanning high-precision and long-horizon behaviors. With only 20 initial demonstrations, Res-HIL outperforms state-of-the-art full-policy human-in-the-loop reinforcement learning and residual fine-tuning without human guidance on every task after ten minutes of online training. Res-HIL improves its pretrained base policies and outperforms imitation policies trained with five times more demonstrations. An ablation study shows that direct residual supervision is critical to performance, while intervention-aware reward shaping substantially improves training efficiency.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ReF-HIL: Shaping the Critic around Human Action Neighborhoods for Efficient Human-in-the-Loop Reinforcement Learning

    Sep 29, 2026Shaoyin Luo, Song Wang, Shibo Xia +5Reinforcement Learning From Human FeedbackHuman-In-The-Loop

  2. Dexterous Robot Manipulation from Human Demonstrations via Contact-Anchored Retargeting and Residual Policy Learning

    Sep 21, 2026Zihao Yang, Chengyuan Liu, Yu Zhou +6Cross-Embodiment RetargetingBimanual Manipulation

  3. RoHIL: Robust Human-in-the-Loop Robotic Reinforcement Learning Against Illumination Variations

    May 19, 2026Shuoqin Zhang, Yixin Xiong, Xiru Gao +4Scalable Robot LearningHuman-In-The-Loop