cs.LGMay 25, 2025

Distributionally Robust Deep Q-Learning

Authors: Chung I Lu, Julian Sester, Aijia Zhang

Organizations: National University of Singapore, Department of Mathematics, 21 Lower Kent Ridge Road, 119077.

Abstract

We propose a novel distributionally robust QQ-learning algorithm for the non-tabular case accounting for continuous state spaces where the state transition of the underlying Markov decision process is subject to model uncertainty. The uncertainty is taken into account by considering the worst-case transition from a ball around a reference probability measure. To determine the optimal policy under the worst-case state transition, we solve the associated non-linear Bellman equation by dualising and regularising the Bellman operator with the Sinkhorn distance, which is then parameterised with deep neural networks. This approach allows us to modify the Deep Q-Network algorithm to optimise for the worst case state transition. We illustrate the tractability and effectiveness of our approach through several applications, including a portfolio optimisation task based on S&{P}~500 data. We also establish convergence guarantees for exact and approximate robust fitted QQ-iteration, decompose the numerical RDQN error into interpretable components, and discuss extensions to compact continuous action sets.

Figures & tables

Explore similar work

CardsList
  1. Finite-Time Convergence of Single-Trajectory Chi-Square Robust Q-Learning With Linear Function Approximation

    Oct 2, 2025Saptarshi Mandal, Yashaswini Murthy, R. SrikantQ-LearningOffline Reinforcement Learning

  2. Mini-Batch Risk-Averse Deep Q-Learning: A Robot Navigation Case Study

    Sep 7, 2026Aayush Patel, Andrzej RuszczyńskiMarkov Decision ProcessesRobot Navigation

  3. Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning

    Jun 8, 2025Yang Xu, Swetha Ganesh, Vaneet AggarwalQ-LearningSoft Actor-Critic