cs.LGSep 30, 2026

Free Everywhere, Exact on Trees: PPO's Dropped Correction Buys Sample Efficiency Under Aggressive Reuse

Authors: Nima H. Siboni

Organizations: Juna.ai, Kastanienallee 32, 10435 Berlin, Germany

Abstract

Common policy improvement methods, including TRPO, PPO, and GRPO, estimate policy improvement under the behavioral policy's state-visitation distribution rather than the improved policy's own. The substitution makes the objective estimable from the behavioral policy's rollouts but adds a bias growing with policy divergence, hence the trust region or clip, and hence no reuse of a batch far off-policy. We show that under history-injective dynamics, where each state is reached by exactly one history, the dropped state-visitation ratio equals the product of per-step policy ratios along the sampled prefix, on every trajectory and not only in expectation. The ratio is therefore restored exactly, from log-probabilities PPO already computes. Autoregressive generation and canonical-order constructive optimization are both history-injective. The exact correction pays importance-sampling variance that grows with the horizon, so we generalize it to a one-parameter family with PPO (α=0α{=}0) and the full correction (α=1α{=}1) as endpoints: a single bias--variance knob. A gradient-level analysis of the unclipped surrogate identifies two channels the correction acts through and three conditions under which it carries signal; an enumerable testbed confirms the conditions' predictions. On hard credit-assignment scheduling tasks, a short corrected warmup with aggressive early sample reuse learns faster than PPO and than the same reuse uncorrected; the marginal gain grows with task difficulty (+0.02+0.02 to +0.09+0.09 learning-curve AUC), and the early win over PPO tracks the prefix bias that reuse incurs. A correction held throughout, or applied where clipping already contains the reuse bias, is null to harmful.

Figures & tables

Appendix figures & tables32 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Reusing Past Samples in Proximal Policy Optimization: When and How Does It Help?

    Oct 1, 2026Alessandro Montenegro, Riccardo Venturelli, Marco Mussi +2Proximal Policy OptimizationOff-Policy Learning

  2. OPTS-TTPO: Enhancing Finite-Sample Policy-Gradient Learning with Tree Search

    Sep 30, 2026Junyu Lu, Shichao Weng, Zhiqiang Wang +10Frictive Policy OptimizationPolicy Gradient