With massively parallel simulation, on-policy Reinforcement Learning methods such as PPO have become standard in many domains. However, learning from scratch is sample-inefficient and fails to exploit the potential existence of a suboptimal expert, such as a heuristic, a model-based controller, or a policy trained on a related task. Such an expert is often available and can guide early training, but its sub-optimality limits final performance. The challenge then becomes balancing expert guidance against learning from rewards. Existing methods set the expert's influence through a blending weight, a schedule, or an evaluation-driven curriculum. Alternatively, they adapt it with additional learned components such as critics over expert actions or auxiliary agents. However, none optimizes it using the same on-policy objective as the policy itself. We propose a method in which the learner and the expert alternate control within each training episode, and the expert's share of control is a single learnable parameter optimized jointly with the policy. The learner benefits from the expert early in training, but its share of control declines as the learner becomes more competent, until eventually vanishing completely. This leaves the learner acting alone and better than the suboptimal expert. We evaluate our method on 34 tasks across two benchmarks, spanning discrete and continuous action spaces, using both learned and model-based experts. Our method improves sample efficiency over guided and unguided baselines while requiring minimal hyperparameter variation. The expert's share decays to zero as the learner improves, vanishing when the expert is no longer useful.
Figures & tables
Figure 1: The expert share is learned from block-level performance. The expert (orange) and learner (blue) alternate control in temporally committed blocks, with λ determining the expert’s share. The difference between expert- and learner-block advantages drives the update of λ : as the learner improves, the expert is progressively withdrawn until λ=0 .
Figure 2: Temporally committed mixture. The horizon is split into blocks of k steps, starting at bj=jk (filled states). At each boundary, a switch zj∼Bernoulli(λ) hands the whole block to the expert π^ (orange, zj=1 ) or to the learner πθ (blue, zj=0 ).
Figure 3: Aggregate results on both benchmarks. Normalized return (min-max per task, averaged over tasks and seeds) and the expert share λ per task.
Figure 4
Figure 5: Frozen Lake with a failing expert ( pslip=0.3 ). 60 evaluation episodes per method on the same map, with the success rate above each panel. Only our learner reliably completes the route.
Figure 6: Hyperparameter ablations. (a) Imitation weight: normalized return vs. cI (Fig. 11 ). (b) Expert price: withdrawal step of λ (top) and relative return (bottom) vs. cE (Fig. 12 ). (c) Block size trade-off: dotted line marks k=100 . Stars mark the maximum. Per-task plots in App. B.4 .
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Task
bar
Ours
JSRL
Kick
BC → PPO
λ 1M
λ 2.5M
λ 5M
PPO
AcrobotSwingup
145
5.8 / 175
9.3 / 156
9.1 / 152
7.7 / 192
8.7 / 185
7.8 / 193
8.7 / 189
14.8 / 189
AcrobotSwingupSparse
41
8.6 / 54
12.3 / 36
– / 27
12.4 / 50
14.0 / 52
14.0 / 50
11.9 / 49
– / 19
BallInCup
730
1.9 / 974
3.3 / 973
4.2 / 968
2.2 / 971
2.9 / 973
3.3 / 974
4.6 / 973
6.6 / 974
CartpoleBalance
665
1.0 / 854
1.3 / 822
1.3 / 834
0.1 / 884
1.3 / 817
1.3 / 814
1.7 / 821
2.5 / 886
CartpoleBalanceSparse
750
1.1 / 987
3.3 / 1000
8.7 / 743
16.5 / 642
4.2 / 996
5.0 / 998
5.0 / 933
6.0 / 944
CartpoleSwingup
566
1.5 / 661
6.6 / 704
2.5 / 572
5.7 / 719
3.3 / 697
4.6 / 690
2.5 / 693
4.2 / 754
Appendix
Table 2: Sample efficiency and final performance per task on DMC. Each cell shows steps to bar (M) / asymptotic return (median over 5 seeds; “–” if not reached within 20M steps). In the median-steps row, runs that do not reach the bar count as 20M. Normalized return is the return divided by the best method’s return on that task, averaged over tasks. Bold : best. Underline : second best. Ties are marked for every tied method and counted for each of them in the last row.
Figure 7: DMControl learning curves per task. Return of the learner acting alone with deterministic actions, for each method.
Figure 8: DMControl expert share per task. Expert share λ of our method over training.
Task
bar
Ours
JSRL
Kick
BC → PPO
λ 1M
λ 2.5M
λ 5M
PPO
KeyCorridorS4R3
0.75
4.7 / 0.98
5.9 / 0.99
16.7 / 0.75
2.2 / 0.94
4.2 / 0.92
4.9 / 1.00
7.4 / 0.99
– / 0.00
KeyCorridorS5R3
0.72
8.5 / 0.96
7.9 / 0.92
17.9 / 0.69
– / 0.00
9.1 / 0.77
9.1 / 0.94
12.4 / 0.90
– / 0.00
KeyCorridorS6R3
0.67
7.8 / 0.90
12.3 / 0.77
15.7 / 0.75
– / 0.00
– / 0.00
– / 0.46
14.8 / 0.77
– / 0.00
DoorKey-16x16
0.70
8.5 / 0.93
10.5 / 0.79
16.7 / 0.69
8.6 / 0.81
9.0 / 0.90
9.6 / 0.90
12.5 / 0.77
– / 0.17
DoorKey-Random-16x16
0.52
10.3 / 0.69
– / 0.15
– / 0.34
– / 0.09
– / 0.21
15.8 / 0.49
14.5 / 0.61
– / 0.03
LavaCrossingS11N5
0.48
– / 0.42
12.5 / 0.42
– / 0.30
4.3 / 0.58
8.7 / 0.56
10.9 / 0.64
18.9 / 0.45
– / 0.06
Appendix
Table 3: Sample efficiency and final performance per task on MiniGrid. Each cell shows steps to bar / final return . Steps to bar is the number of environment steps (in millions) needed to reach 75% of the best final return on that task (median over 5 seeds; “–” if not reached within 20M steps). Final return is the deterministic return after 20M steps (mean over 5 seeds). In the median-steps row, runs that do not reach the bar count as 20M. Normalized return is the final return divided by the best method’s final return on that task, averaged over tasks. Bold : best. Underline : second best. Ties are marked for every tied method and counted for each of them in the last row.
Figure 9: MiniGrid learning curves per task. Return of the learner acting alone with deterministic actions, for each method. The dotted line is the expert’s return.
Figure 10: MiniGrid expert share per task. Expert share λ of our method over training.
Method
Success rate
PPO
0.000
BC → PPO
0.000
Kickstarting
0.037
JSRL
0.276
Ours
0.985
Appendix
Table 4: Final success rate on FrozenLake ( pslip=0.3 ), mean over 5 seeds.
Figure 11: Imitation weight per task. Return over training on six DMControl tasks for each cI . With cI∈{0,0.01} , the learner stays near zero return on the three Humanoid tasks. With cI=1 , it stays near the expert’s return on FingerSpin and ends lower on the Humanoid tasks. On BallInCup and WalkerWalk, cI has little effect.
Figure 12: Expert price per task. Expert share λ (top) and return (bottom) over training on three DMControl tasks for each cE . With cE=0 , λ stays above zero until the end of training on WalkerWalk and CheetahRun, and reaches zero on FingerSpin after about 10M steps. With cE=3 , λ reaches zero within 1M steps on all three tasks.
Parameter
Value
Shared PPO backbone
N parallel environments
2048
Rollout length T
30
Minibatches / epochs
32 / 16
γ / λGAE
0.995 / 0.95
Clip ϵ
0.2
Appendix
Table 5: Hyperparameters. The PPO backbone is identical for all methods; only the guidance mechanism differs. All methods use the same frozen guide πE with no per-task tuning.
Parameter
Value
Shared PPO backbone
N parallel environments
512
Rollout length T
64
Minibatches / epochs
8 / 2
Learning rate
10−3
Observation normalisation
no
Appendix
Table 6: MiniGrid hyperparameters. Only differences from the DMControl settings in Table 5 are reported.
Task
Guide πE
ceiling
πE/ceil
AcrobotSwingup
139.1
193
0.72
AcrobotSwingupSparse
15.7
54
0.29
BallInCup
844.0
974
0.87
CartpoleBalance
724.9
886
0.82
CartpoleBalanceSparse
256.2
1000
0.26
CartpoleSwingup
693.1
754
0.92
Appendix
Table 7: Guide quality on DMControl. The ceiling is the best asymptotic return of any method on that task (Table 2 ). Guides are deliberately sub-optimal (mean 0.60× the ceiling, range 0.26 – 0.97 ).
Task
Guide πE
ceiling
πE/ceil
KeyCorridorS4R3
0.546
1.000
0.55
KeyCorridorS5R3
0.703
1.000
0.70
KeyCorridorS6R3
0.825
1.000
0.83
DoorKey-16x16
0.776
1.000
0.78
DoorKey-Random-16x16
0.857
1.000
0.86
LavaCrossingS11N5
0.154
1.000
0.15
Appendix
Table 8: Guide quality on MiniGrid. Guides are deliberately sub-optimal (mean 0.73× the attainable performance).