A Closed-Loop Non-Asymptotic Convergence Analysis of PPO with Learned Critics and Clipping
Organizations: School of Artificial Intelligence and Data Science, University of Science and Technology of China · Department of Computer Science, University of Hong Kong
Abstract
Despite its widespread use, Proximal Policy Optimization with clipping (PPO-Clip) remains difficult to tune, and the interactions among critic learning, clipping, and rollout reuse remain incompletely understood. We develop a \emph{non-asymptotic} analysis of PPO-Clip as a \emph{closed-loop actor--critic} system. It captures actor--critic coupling, nonsmooth probability-ratio clipping, finite-batch reuse, and predictable early stopping under explicit coverage and critic regularity assumptions, using raw GAE and Monte Carlo critic targets. Our synchronous and asynchronous guarantees jointly characterize policy stationarity and the tracking accuracy of the learned critic, with explicit dependence on algorithmic parameters. A sufficient coupling condition gives optimization, critic tracking, clipping, and finite-batch errors a common amplification bound. The asynchronous result also requires a delay-dependent critic stepsize restriction; violating these conditions does not establish divergence. For finite layered MDPs with tabular critics, a uniform bound on the actual clipped-gradient class replaces complete-trajectory counting. A verified growing-horizon family has polynomial sample complexity, and a two-time-scale schedule gives stationarity and critic-tracking bounds with explicit fresh-rollout accounting. These results together advance our understanding about PPO and provide theoretical guidance in tuning.
Figures & tables
| Work | Actor update | Critic treatment | Statistical control | Conclusion |
|---|---|---|---|---|
| Jin et al. (2023) | Clipped surrogate gradients | Estimated advantages; no critic recursion | Sampling/advantage bias retained as | Stationarity up to bias |
| Huang et al. (2024) | Entropic mirror descent; neural policy regression | Neural action-value TD | Evaluation/improvement errors controlled by width and inner iterations | Global optimality for analyzed scheme |
| Wu et al. (2020) | Online score update with TD error | Simultaneous linear TD(0) | Markovian data; uniform mixing; critic tracking | Stationarity up to approximation error |
| Doering et al. (2026) | Symmetric clipped-gradient proxy | Bounded advantage bias assumed | Finite-buffer reuse with random reshuffling | Epoch-start stationarity with error floors |
| This work | Interleaved clipping; recomputed raw GAE | Tabular MC regression on stored returns | Uniform adaptive actor/critic batch bounds; population KL control | Joint stationarity and critic parameter tracking |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Critic | Clipping | Final return | Critic MSE | ||
|---|---|---|---|---|---|
| Slow | 16 | Enabled | 6.03452 | 0.935176 | 0 |
| Slow | 16 | Disabled | 6.03452 | 0.935176 | 0 |
| Slow | 64 | Enabled | 6.01705 | 0.923023 | 0.00204463 |
| Slow | 64 | Disabled | 6.03721 | 0.972968 | 0 |
| Fast | 16 | Enabled | 6.02924 | 0.059509 | 0 |
| Fast | 16 | Disabled | 6.02924 | 0.059509 | 0 |