Imitation learning enables robots to acquire manipulation skills from demonstrations, but the resulting policies can fail outside the training data, while collecting more demonstrations requires substantial human effort. Human-in-the-loop reinforcement learning uses corrective feedback during online training, but typically learns the complete task policy rather than refining a pretrained imitation policy. We introduce Res-HIL, a human-in-the-loop residual reinforcement learning framework that learns corrective actions on top of a frozen imitation policy. Each human intervention provides two complementary learning signals: direct supervision of the residual policy and reward shaping of preceding autonomous behavior. Res-HIL combines these signals with zero initialization of the residual policy to stabilize and accelerate online learning. We evaluate Res-HIL on five contact-rich manipulation tasks spanning high-precision and long-horizon behaviors. With only 20 initial demonstrations, Res-HIL outperforms state-of-the-art full-policy human-in-the-loop reinforcement learning and residual fine-tuning without human guidance on every task after ten minutes of online training. Res-HIL improves its pretrained base policies and outperforms imitation policies trained with five times more demonstrations. An ablation study shows that direct residual supervision is critical to performance, while intervention-aware reward shaping substantially improves training efficiency.
Figures & tables
Fig. 1 : Illustration of the cable task. The robot must route the grasped cable around the cylinder (1) and insert its end into the socket (2), resulting in the successful final configuration shown on the image.
Fig. 2 : Overview of the proposed Res-HIL framework. During environment interaction, the frozen base policy and the learned residual policy jointly determine the executed action, while human interventions provide corrective actions and are incorporated into reward shaping. All transitions are stored in the replay buffer, with intervention transitions additionally stored in the intervention buffer. In parallel, the learning loop continuously samples training batches to update the critics and residual actor, followed by the corresponding target networks.
Method
Peg-in-Hole
Peg-in-Hole
Vent
Cable
Cable
Easy
Hard
2 Cameras
3 Cameras
HIL-SERL
90
30
10
0
0
ResFiT (100 demos)
0
0
0
0
0
Naive Res-HIL
12
0
0
0
0
Res-HIL
100
64
50
66
92
TABLE I : Success rates after 10 minutes of online training. All values are reported as percentages.
Method
Peg-in-Hole
Peg-in-Hole
Vent
Cable
Cable
Easy
Hard
2 Cameras
3 Cameras
ACT (20 demos)
36
28
44
46
70
ACT (100 demos)
84
60
80
80
92
HIL-SERL
100
94
84 ∗
0
0
ResFiT (100 demos)
4
0
0
0
0
Naive Res-HIL
24
12
12
0
0
TABLE II : Final success rates. ACT policies are evaluated without online fine-tuning. All values are reported as percentages.
Method
Peg-in-Hole
Peg-in-Hole
Vent
Cable
Cable
Average
Easy
Hard
2 Cameras
3 Cameras
ACT (20 demos)
6.37
8.17
6.79
7.40
9.90
7.73
ACT (100 demos)
2.88
5.70
6.00
6.64
7.95
5.83
HIL-SERL
2.21
2.65
2.04
–
–
–
Res-HIL
2.67
3.40
5.98
7.01
8.26
5.46
TABLE III : Average cycle time per successful episode across tasks in seconds. Lower values indicate faster task completion.
Variant
Success (%)
Time
Full Res-HIL
100
8m
w/o zero initialization
100
16m
w/o reward shaping
92
30m ∗
w/o BC loss
20
30m ∗
TABLE IV : Component ablation on easy Peg-in-Hole task.
Human-in-the-loop reinforcement learning (HIL-RL) offers a promising route to efficient training of robotic manipulation policies by combining autonomous learning with human demonstrations and online corrections. However, insufficient use of successful human experience in value learning prolongs costly real-world training, while persistent imitation penalties can limit value-driven policy improvement. To address these limitations, we propose ReF-HIL, an efficient HIL-RL framework that uses human guidance to accelerate the learning process. Human-Reference-Guided Value Shaping learns an independent value reference from successful human experience to guide online value learning, while incorporating local corrective feedback. A Human Action Fence defines a learned human-action neighborhood, allowing value-driven optimization for better performance without imitation penalties inside while constraining policy and value updates outside. Experiments on five diverse and challenging real-world manipulation tasks demonstrate improved overall learning efficiency and higher success rates compared with the evaluated baselines. Specifically, ReF-HIL reaches 90% autonomous success in only 18-63 minutes of active training and achieves final success rates of 91.7-100%. These results highlight the potential of human-guided reinforcement learning to acquire reliable manipulation skills efficiently in the real world. Project website: https://anonymous.4open.science/w/ReF-HIL-7762/
Shaoyin Luo, Song Wang, Shibo Xia +5
Department of Mechanical Engineering, Tsinghua University, Beijing 100084, China
Learning dexterous manipulation from demonstrations is bottlenecked by data: the contact forces that determine whether a grasp succeeds are absent from every scalable source of human demonstrations. This paper builds on two observations. First, what survives the change from a human hand to a robot hand is the contact structure of a demonstration - which finger regions touch which object locations, and in what order - rather than its joint motion. Second, physical consistency need not be engineered per task: a single residual reinforcement learning (RL) policy, trained once across diverse demonstrations, can repair kinematic recordings into physically consistent, contact-annotated trajectories, and the same residual formulation restores dynamic feasibility after retargeting. These observations yield a three-stage pipeline that converts human motion-capture recordings into dexterous robot policies with no real-robot training data: physics refinement with a simulated MANO hand recovers contacts and forces, contact-anchored retargeting transfers the demonstrated contact structure through an objective independent of hand morphology, and residual policy learning adapts the result to robot actuation. The pipeline reconstructs 25,454 single-hand trajectories (success 7.3% -> 59.3%) and 25 dual-hand tasks (16.0% -> 62.4%) with one shared policy per setting, transfers one human dataset to four morphologically distinct robot hands (+62.4 pp), and executes four contact-rich bimanual tasks on physical hardware with zero real-robot training data.
Zihao Yang, Chengyuan Liu, Yu Zhou +6
DexGEM Lab · Shanghai Jiao Tong University · Tongji University +1
Human-in-the-loop reinforcement learning systems achieve near-perfect success on the workstation where they are trained, but collapse when the same robot is moved to a workstation a few meters away due to shifts in the visual input distribution caused by new lamp positions and window light. Re-collecting demonstrations and re-running HIL on every workstation is incompatible with deployment, and naively fine-tuning on shifted-light data triggers catastrophic forgetting of the source workstation. To close this cross-domain gap, we present RoHIL, an offline fine-tuning framework that uses no extra real-robot interaction. RoHIL combines (i) a world-model-based image relighter that re-synthesises the visual stream of source-workstation trajectories under multiple virtual HDRI environments, leaving actions and rewards real; (ii) Illumination-Retention Replay (IRR), a data-level anti-forgetting mechanism that interleaves relit adaptation transitions with original-light retention transitions to preserve source-workstation Bellman coverage; and (iii) an anchored Bellman-actor regulariser that constrains representation and policy drift from the original source-workstation policy. Across four real-robot manipulation tasks under significant cross-workstation illumination variations, RoHIL substantially improves shifted-light performance where standard HIL-RL collapses, while preserving source-workstation performance, eliminating the need to re-collect data and retrain for every new workstation and environment. Project page: https://anonymous4365.github.io/RoHIL/