Imitation learning enables robots to acquire manipulation skills from demonstrations, but the resulting policies can fail outside the training data, while collecting more demonstrations requires substantial human effort. Human-in-the-loop reinforcement learning uses corrective feedback during online training, but typically learns the complete task policy rather than refining a pretrained imitation policy. We introduce Res-HIL, a human-in-the-loop residual reinforcement learning framework that learns corrective actions on top of a frozen imitation policy. Each human intervention provides two complementary learning signals: direct supervision of the residual policy and reward shaping of preceding autonomous behavior. Res-HIL combines these signals with zero initialization of the residual policy to stabilize and accelerate online learning. We evaluate Res-HIL on five contact-rich manipulation tasks spanning high-precision and long-horizon behaviors. With only 20 initial demonstrations, Res-HIL outperforms state-of-the-art full-policy human-in-the-loop reinforcement learning and residual fine-tuning without human guidance on every task after ten minutes of online training. Res-HIL improves its pretrained base policies and outperforms imitation policies trained with five times more demonstrations. An ablation study shows that direct residual supervision is critical to performance, while intervention-aware reward shaping substantially improves training efficiency.
Figures & tables
Fig. 1 : Illustration of the cable task. The robot must route the grasped cable around the cylinder (1) and insert its end into the socket (2), resulting in the successful final configuration shown on the image.
Fig. 2 : Overview of the proposed Res-HIL framework. During environment interaction, the frozen base policy and the learned residual policy jointly determine the executed action, while human interventions provide corrective actions and are incorporated into reward shaping. All transitions are stored in the replay buffer, with intervention transitions additionally stored in the intervention buffer. In parallel, the learning loop continuously samples training batches to update the critics and residual actor, followed by the corresponding target networks.
Method
Peg-in-Hole
Peg-in-Hole
Vent
Cable
Cable
Easy
Hard
2 Cameras
3 Cameras
HIL-SERL
90
30
10
0
0
ResFiT (100 demos)
0
0
0
0
0
Naive Res-HIL
12
0
0
0
0
Res-HIL
100
64
50
66
92
TABLE I : Success rates after 10 minutes of online training. All values are reported as percentages.
Method
Peg-in-Hole
Peg-in-Hole
Vent
Cable
Cable
Easy
Hard
2 Cameras
3 Cameras
ACT (20 demos)
36
28
44
46
70
ACT (100 demos)
84
60
80
80
92
HIL-SERL
100
94
84 ∗
0
0
ResFiT (100 demos)
4
0
0
0
0
Naive Res-HIL
24
12
12
0
0
TABLE II : Final success rates. ACT policies are evaluated without online fine-tuning. All values are reported as percentages.
Method
Peg-in-Hole
Peg-in-Hole
Vent
Cable
Cable
Average
Easy
Hard
2 Cameras
3 Cameras
ACT (20 demos)
6.37
8.17
6.79
7.40
9.90
7.73
ACT (100 demos)
2.88
5.70
6.00
6.64
7.95
5.83
HIL-SERL
2.21
2.65
2.04
–
–
–
Res-HIL
2.67
3.40
5.98
7.01
8.26
5.46
TABLE III : Average cycle time per successful episode across tasks in seconds. Lower values indicate faster task completion.
Variant
Success (%)
Time
Full Res-HIL
100
8m
w/o zero initialization
100
16m
w/o reward shaping
92
30m ∗
w/o BC loss
20
30m ∗
TABLE IV : Component ablation on easy Peg-in-Hole task.