Authors: Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv
Organizations: Applied and Theoretical Aspects of Robot Intelligence (ATARI) Lab, Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich
Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unreliable policy updates. We propose Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm that performs policy improvement using critic values without differentiating the critic with respect to actions. FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and Kullback-Leibler (KL) divergence. We then fit the actor to this target by minimizing a forward-KL objective, estimated using self-normalized importance sampling (SNIS) with actions drawn from the rollout policy. By limiting the target distribution's deviation from the rollout policy, the KL regularization helps keep these importance weights well behaved. In contrast to reverse-KL objectives, which can favor a subset of the target distribution's modes, the forward-KL objective encourages coverage of multiple high-value modes and thereby promotes exploration. Experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains. Computational benchmarks also demonstrate faster actor updates than Relative Entropy Pathwise Policy Optimization (REPPO).
Figures & tables
Figure 1: Actor directions and local losses. Top: green/red arrows point toward/away from the cost minimum (star); lengths are illustrative, not convergence guarantees. Bottom: frozen-reference losses at A–C, zeroed at the reference; KL losses are scaled by T . Details: Appendix A.1 .
Figure 2: Update cost on an A100. Medians and ranges over three repetitions with 131,072 rollout states. Actor time includes cache construction; total updates include critic training. PyTorch allocation excludes simulation and unused reservations.
Figure 3: Ablations. Left: G1 candidate count and temperature at 50.07M steps (eight seeds for K=2,16 and τ=0 ; ten for K=4,8 ). Middle: cached versus uncached G1 updates ( K=8 ). Right: state-value versus Q critics on AcrobotSwingupSparse (sparse rewards, K=16 ). Means with one SEM (left, middle) or one SD (right); panels use separate controls. The state-value variant uses 17× more simulator transitions per rollout step.
Figure 4: Benchmark learning curves. IQM with equal environment weights and equal seed weights within environments. Bands: pointwise 95% confidence intervals from 10,000 within-environment bootstrap resamples. G1/T1 FERPO uses K=16 , ESS target 0.65 , and eight seeds per task. Details: Appendix A.3 .
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Frozen-reference losses at B. Left: REPPO’s active expected-cost/entropy branch (purple) and KL-only branch (orange); the pale dashed curve extends the inactive cost branch. Open and filled endpoints show the loss switch. Middle: forward KL from the fixed FERPO target to the candidate Gaussian. Right: reverse KL to that same target. Both KL curves are scaled by T=0.15 ; this positive scaling preserves directions. Vertical dotted lines and black dots mark the reference μ0=−0.6 .
Figure 6: Wider proposals recover the population fitting loss. Losses at B with the same reference policy and target throughout. Teal: SNIS median and pointwise 5th–95th percentiles across 5,000 batches; dotted black: population FERPO; purple: analytic REPPO. Percentages give the frequency of rightward SNIS directions. Shading shows batch variation, not confidence intervals for the mean. Losses are zeroed at the reference, and forward losses are scaled by T .
Proposal σr
K
Median m
5th–95th percentiles
Rightward
0.4
64
−0.916
[−0.963,−0.416]
6.6%
0.8
64
−0.498
[−0.863,−0.066]
67.1%
1.2
64
−0.472
[−0.753,−0.091]
75.5%
0.4
1,024
−0.824
[−0.923,0.275]
26.4%
0.8
1,024
−0.468
[−0.565,−0.369]
98.9%
1.2
1,024
−0.466
[−0.546,−0.381]
99.9%
Appendix
Table 1: SNIS target-mean estimates at B. The population target mean is −0.465943… and rightward movement means m>−0.6 . Intervals are the empirical 5th–95th percentiles across 5,000 batches, not confidence intervals. Policy standard deviation remains 0.4 in every row.
Process peak
PyTorch allocated peak
PyTorch reserved
Method
Rollout
Update
Rollout
Update
Rollout
REPPO
1,438
1,408
601
422
614
FERPO
1,760
1,728
600
658
936
Cached FERPO
1,806
1,774
600
723
982
Appendix
Table 2: Memory across rollout and update phases (MiB). Peaks are identical across the two matched seeds. Reserved memory includes active tensors.
Figure 7: Temperature learning curves. Adaptive temperature versus no entropy regularization at K=16 . Both curves show seed means with one-SEM bands. Both curves use eight completed seeds at every evaluation; the adaptive control uses corrected temperature learning rate 0.0003 . Final returns are 34.81±0.24 with adaptive temperature and 30.32±0.27 without entropy regularization. The final evaluations also appear in Figure 3 (left).
Figure 8: Final return versus critic width on G1. Points show seed means and error bars show one SEM. The open square marks the single-seed REPPO reference at width 1,024, for which no error bar is reported. Both methods lose substantial return at width 62; FERPO reaches 6.29 and REPPO 4.19. These endpoints do not identify the source of critic approximation error.
Figure 9: Cached versus uncached Q against elapsed time. G1 deterministic evaluation return for seeds 3–5 per method on the same Vega A100 node. Both use K=8 and eight critic epochs on the full rollout. Bold curves and one-SEM bands use linear interpolation at common times; faint curves show individual runs. The arrow marks mean completion times at 50,069,504 interactions: 55.7 versus 68.4 minutes. At 50 minutes, interpolated mean returns are 33.55 cached and 32.86 uncached. Figure 3 (middle) shows all five seeds against environment interactions.
Figure 10: Final return versus controller target. HopperHop (top) and AcrobotSwingupSparse (bottom), with ESS targets (left) and policy-KL bounds (right). Points show mean final return at 50.07M interactions with one SEM over 12 seeds.
Figure 11: Achieved ESS under both controllers. Per-iteration geometric mean of normalized SNIS ESS under direct ESS control (left) and policy-KL control (right). Curves show unsmoothed means with one-SEM bands across 12 seeds. Dotted lines mark ESS targets; right-panel legends give policy-KL bounds in nats.
Group
Method
Selected configuration or source
DMC
REPPO
Official published trials
DMC
PPO
ppo_4
DMC
PPO (Brax)
ppo_brax
DMC
FastTD3
fasttd3_small
G1/T1
REPPO
Official published trials
ManiSkill
REPPO
Official published trials
Appendix
Table 3: Baseline sources, fixed across tasks within each group. The archive contains no PPO or FastTD3 G1/T1 curves, and no FastTD3 ManiSkill curves.
Suite
Method A − method B
IQM difference
Adjusted interval
MuJoCo DMC
FERPO − REPPO
6.67
[−6.89,20.65]
MuJoCo DMC
FERPO − FastTD3
7.10
[−11.82,30.90]
MuJoCo DMC
FERPO − PPO
587.89
[570.19,604.70]
MuJoCo DMC
FERPO − PPO (Brax)
354.86
[307.06,398.14]
MuJoCo DMC
REPPO − FastTD3
0.44
[−18.30,23.87]
MuJoCo DMC
REPPO − PPO
581.22
[563.44,597.76]
Appendix
Table 4: Pairwise differences in final suite IQM. Intervals use 100,000 seed bootstrap resamples and a Bonferroni correction across all fourteen comparisons (95% family-wise target). MuJoCo differences are in return units; ManiSkill differences are in percentage points (pp). Positive values favor method A. The G1/T1 comparison uses eight FERPO seeds per task.
Figure 12: DMC learning curves (1/2). The first 12 of 22 tasks. All methods show IQM with pointwise 95% bootstrap confidence intervals. Legends report contributing seeds n ; ranges indicate changes across the displayed evaluations. Curves retain their source evaluation schedules. Baseline configurations are fixed across all DMC tasks as specified in Table 3 .
Figure 13: DMC learning curves (2/2). The remaining ten tasks, under the same protocol as Figure 12 . All methods show IQM with pointwise 95% bootstrap confidence intervals. Walker markers denote the available checkpoint evaluations from 3.01M to 50.99M steps, with ten seeds per task. Legends report contributing seed counts.
Figure 14: All four G1/T1 locomotion tasks. FERPO and official REPPO trials show IQM with pointwise 95% bootstrap confidence intervals. FERPO uses K=16 , normalized ESS target 0.65 , and eight completed seeds per task; legends report each method’s seed count. Results use the recorded evaluation protocols described in Appendix A.3.7 .
Figure 15: All eight ManiSkill tasks. All methods show success-rate IQM with pointwise 95% bootstrap confidence intervals. FERPO uses 20 seeds per task except G1 transport box (ten). PPO uses one configuration across all eight tasks; REPPO uses the official published trial CSVs.
We present Reflow-regularized Flow Matching Policy Gradients (ReFPO), a simple online RL method that adds explicit Reflow regularization to FPO for efficient flow-based control. We uncover a key structural property: the gradient updates in Flow Matching Policy Gradients (FPO) can be interpreted as an implicit advantage-weighted Reflow process, providing a new geometric perspective on flow-based policy gradients. Building on this insight, ReFPO introduces an explicit geometric regularizer that can be implemented with a single line of code change without incurring additional computational overhead or auxiliary distillation stages. By synergizing advantage-guided updates with path rectification, our method reduces CFM proxy-ratio spikes, stabilizes PPO-style training, and enables high-fidelity one-step inference that often matches or exceeds multi-step performance. We experimentally demonstrate that ReFPO improves average performance and discretization robustness across GridWorld, MuJoCo Playground, and high-dimensional Humanoid Control tasks, providing a scalable and stable approach for generative policies in complex physical simulations.
In this paper, we study the role of the critic in actor--critic for entropy-regularized, finite, discounted environments. We establish that, when the critic is exact, using the latter as a baseline is a variance-reduction method in a strong sense. In this case, actor--critic with stochastic gradients matches the sample complexity of deterministic policy gradient, reaching an ε-optimal regularized value with O~(log(1/ε)) samples. In practice, the critic is learned alongside the actor: the variance of the actor update is then influenced by the critic's variance and bias. Specifically, when the critic has a sufficiently small error, the variance reduction and rapid convergence are preserved. This suggests to learn the critic first, keeping it up to date after each actor update, underscoring the crucial role of accurate critic estimation in actor--critic methods.
Safwan Labbi, Paul Mangold, Daniil Tiapkin +1
CMAP, CNRS, École Polytechnique, Institut Polytechnique de Paris, 91120 Palaiseau, France · Université Paris-Saclay, CNRS, LMO, 91405, Orsay, France · Google DeepMind +2
Proximal Policy Optimization (PPO) is the standard policy-gradient algorithm for on-policy reinforcement learning. The literature presents it in two forms, a clipped surrogate that bounds the importance ratio between successive policies and a Kullback-Leibler penalty between them. These forms are treated as separate algorithms with their own gradients, their own hyperparameters, and their own reference implementations, and a sizeable body of empirical work compares them. We show that the gradient of the clipped surrogate is reproduced exactly by a Kullback-Leibler surrogate whose coefficient varies per sample, with closed-form dependence on the importance ratio and the advantage. The identity holds at every minibatch step and across the entire inner loop, and on five MuJoCo continuous-control benchmarks the two losses produce indistinguishable training curves. The reformulation exposes a structural feature of the clipped surrogate that the min notation hides. PPO-Clip's implicit per-sample penalty is a step function at the boundary of the trust region, and the shape of this coefficient is the natural design axis for generalising the algorithm. We sketch the resulting follow-up directions in the discussion.