cs.ROJul 28, 2026

P3: Probabilistic Policy Propagation for Stable VAE-Based Robot Learning

Authors: Liyun YanJianming MaYang ZhangShengcheng FuZhanxiang CaoKeqi ZhuYizhi ChenYue Gao

Organizations: Shanghai Jiao Tong University, Shanghai, China · Shanghai Innovation Institute, Shanghai, China · Tongji University, Shanghai, China · Zhejiang University, Hangzhou, China

Abstract

Variational Autoencoders are widely used to encode high-dimensional and noisy observations in robotics. However, their stochastic latent creates a mismatch with Proximal Policy Optimization (PPO): an effective policy marginalizes over the latent distribution, whereas former implementations estimate its probability ratio and KL divergence using only one latent sample. We identify a fundamental but overlooked theoretical cause: naive single-sample approximations in stochastic latent space induce significant variance and bias in the surrogate loss. To address this, we introduce P^3 (Probabilistic Policy Propagation), a distribution-aware optimization framework for VAE-based policies. P3P^3 couples moment-based probabilistic method for stable and efficient learning with sampling-based calibration for robust policy behavior under latent uncertainty. In our experiments, P^3 boosts data efficiency from 64.6% to >96%, reduces convergence steps by >20%. Furthermore, P^3 is evaluated on challenging humanoid parkour tasks and shows an effective foundation for VAE-based PPO. Code is available at https://github.com/ylyem9x/P3_Open.

Explore similar work

CardsList
  1. PAC-DP: PAC-Bayesian Diffusion Policy Learning

    Jul 27, 2026Mohammad Hasan Yeganegi, Dian Yu, Andrea Del Prete +2Robot PoliciesBayesian