cs.LGOct 5, 2026

Adaptive Expert Guidance for Efficient On-Policy Reinforcement Learning

Authors: Daniele Affinita, Ming Xu, Rudolf Reiter, Davide Scaramuzza, Pascal Fua

Organizations: EPFL, Switzerland · University of Zurich, Switzerland

Abstract

With massively parallel simulation, on-policy Reinforcement Learning methods such as PPO have become standard in many domains. However, learning from scratch is sample-inefficient and fails to exploit the potential existence of a suboptimal expert, such as a heuristic, a model-based controller, or a policy trained on a related task. Such an expert is often available and can guide early training, but its sub-optimality limits final performance. The challenge then becomes balancing expert guidance against learning from rewards. Existing methods set the expert's influence through a blending weight, a schedule, or an evaluation-driven curriculum. Alternatively, they adapt it with additional learned components such as critics over expert actions or auxiliary agents. However, none optimizes it using the same on-policy objective as the policy itself. We propose a method in which the learner and the expert alternate control within each training episode, and the expert's share of control is a single learnable parameter optimized jointly with the policy. The learner benefits from the expert early in training, but its share of control declines as the learner becomes more competent, until eventually vanishing completely. This leaves the learner acting alone and better than the suboptimal expert. We evaluate our method on 34 tasks across two benchmarks, spanning discrete and continuous action spaces, using both learned and model-based experts. Our method improves sample efficiency over guided and unguided baselines while requiring minimal hyperparameter variation. The expert's share decays to zero as the learner improves, vanishing when the expert is no longer useful.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Expert Behavior Prior Reinforcement Learning

    Jul 23, 2026Gong Gao, Weidong Zhao, Xianhui Liu +1Model-Based Reinforcement LearningOffline Reinforcement Learning

  2. When (and How) to Trust the Expert: Diagnosing Query-Time Expert-Guided Reinforcement Learning

    May 9, 2026Yann Berthelot, Philippe Preux, Riad AkrourExpertsAI Trustworthiness

  3. MInTRL: Off-policy Intervention can boost On-policy RL

    Sep 14, 2026Mingyu Chen, Yefan Tao, Gerald Friedland +2Off-Policy Learning