cs.LGMay 23, 2024

On the Approximation and Convergence of Distributional Policy Gradient Algorithms for Risk-Sensitive Reinforcement Learning

Authors: Xian Yu, Minheng Xiao, Lei Ying

Organizations: Department of Integrated Systems Engineering, The Ohio State University, Columbus, OH, USA · Department of Electrical Engineering and Computer Science, University of Michigan, Ann Arbor, MI, USA

Abstract

Risk-sensitive reinforcement learning (RL) is crucial for maintaining reliable performance in high-stakes applications. While traditional RL methods aim to learn a point estimate of the random cumulative cost, distributional RL seeks to estimate the entire distribution of it, leading to a unified framework for handling different risk measures. However, developing policy gradient methods for risk-sensitive distributional RL is inherently more complex as it often involves finding the gradient of a probability measure. This paper introduces a new distributional policy gradient framework for risk-sensitive RL, where we derive an analytical gradient of the probability measure of the cumulative cost. For practical implementation, we further design a categorical distributional policy gradient algorithm (CDPG) that approximates arbitrary distributions using a categorical family supported on fixed points. Using Conditional Value-at-Risk (CVaR) as the objective, we prove that the proposed CDPG converges to stationary points and establish its iteration complexities under inexact policy evaluation. Through experiments in a stochastic Cliffwalk environment, we demonstrate the effectiveness of the proposed algorithm and highlight the benefits of incorporating risk sensitivity into distributional RL.

Figures & tables

Explore similar work

CardsList
  1. Reinforcement Learning under State and Outcome Uncertainty: A Foundational Distributional Perspective

    Sep 21, 2026Larry Preuett, Qiuyi Zhang, Muhammad Aurangzeb AhmadPartially Observable Markov Decision Process

  2. Actor-Critic Algorithm for Dynamic Expectile and CVaR

    May 8, 2026Yudong Luo, Erick DelageConditional-Value-At-RiskSoft Actor-Critic