On the Approximation and Convergence of Distributional Policy Gradient Algorithms for Risk-Sensitive Reinforcement Learning
Authors: Xian Yu, Minheng Xiao, Lei Ying
Organizations: Department of Integrated Systems Engineering, The Ohio State University, Columbus, OH, USA · Department of Electrical Engineering and Computer Science, University of Michigan, Ann Arbor, MI, USA
Risk-sensitive reinforcement learning (RL) is crucial for maintaining reliable performance in high-stakes applications. While traditional RL methods aim to learn a point estimate of the random cumulative cost, distributional RL seeks to estimate the entire distribution of it, leading to a unified framework for handling different risk measures. However, developing policy gradient methods for risk-sensitive distributional RL is inherently more complex as it often involves finding the gradient of a probability measure. This paper introduces a new distributional policy gradient framework for risk-sensitive RL, where we derive an analytical gradient of the probability measure of the cumulative cost. For practical implementation, we further design a categorical distributional policy gradient algorithm (CDPG) that approximates arbitrary distributions using a categorical family supported on fixed points. Using Conditional Value-at-Risk (CVaR) as the objective, we prove that the proposed CDPG converges to stationary points and establish its iteration complexities under inexact policy evaluation. Through experiments in a stochastic Cliffwalk environment, we demonstrate the effectiveness of the proposed algorithm and highlight the benefits of incorporating risk sensitivity into distributional RL.
Figures & tables
Fig. 1: Illustration of the stochastic Cliffwalk environment.