cs.LGOct 1, 2026

FERPO: Forward Entropy-Regularized Policy Optimization

Authors: Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv

Organizations: Applied and Theoretical Aspects of Robot Intelligence (ATARI) Lab, Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich

Abstract

Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unreliable policy updates. We propose Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm that performs policy improvement using critic values without differentiating the critic with respect to actions. FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and Kullback-Leibler (KL) divergence. We then fit the actor to this target by minimizing a forward-KL objective, estimated using self-normalized importance sampling (SNIS) with actions drawn from the rollout policy. By limiting the target distribution's deviation from the rollout policy, the KL regularization helps keep these importance weights well behaved. In contrast to reverse-KL objectives, which can favor a subset of the target distribution's modes, the forward-KL objective encourages coverage of multiple high-value modes and thereby promotes exploration. Experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains. Computational benchmarks also demonstrate faster actor updates than Relative Entropy Pathwise Policy Optimization (REPPO).

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ReFPO: Reflow Regularization for Flow Matching Policy Gradients

    Jun 19, 2026Ge Wang, Yibo Peng, Fan Feng +10Flow PoliciesOffline Reinforcement Learning

  2. Refined Analysis of Entropy-Regularized Actor-Critic

    May 23, 2026Safwan Labbi, Paul Mangold, Daniil Tiapkin +1Entropy Regularized Reinforcement LearningSoft Actor-Critic

  3. KLip-PPO: A per-sample KL perspective on PPO-Clip

    Jun 22, 2026Riccardo Colletti, Robin HolzingerProximal Policy OptimizationKullback-Leibler Divergence