math.OCOct 7, 2026

Boundary-aware Reinforcement Learning for Hypercube State Spaces via Deterministic Policy Gradient

Authors: Lijun Bo, Yijie Huang, Chenhao Lu

Abstract

We develop a continuous-time deterministic policy gradient framework for reinforcement learning with reflected state dynamics, where the state process is governed by a controlled reflected stochastic differential equation on a hypercube. Under suitable regularity assumptions, we establish the connection between the value function and the Neumann Bellman equation, introduce an advantage-rate function that yields a deterministic policy gradient formula, and prove the martingale characterization theorem. Motivated by these theoretical results, we propose a continuous-time deep deterministic policy gradient algorithm for reflected stochastic systems, in which the Neumann boundary condition is imposed via either soft penalization or hard architectural constraint. We further quantify the discrepancy between the ideal continuous-time dynamics and the discretely sampled exploratory dynamics executed in practice, showing that the error decays as the time grid is refined and exploration noise vanishes. Our experiments on reservoir control problems illustrate the effectiveness of the RL framework, highlighting that boundary-aware methods substantially reduce Neumann boundary residuals and enhance learning stability.

Figures & tables

Explore similar work

CardsList
  1. From Ticks to Flows: Dynamics of Neural Reinforcement Learning in Continuous Environments

    Jun 2, 2026Saket Tiwari, Tejas Kotwal, George KonidarisSoft Actor-CriticOffline Reinforcement Learning

  2. Policy Gradient for Continuous-Time Robust Markov Decision Processes

    Jun 3, 2026Tanya Veeravalli, David M. Bossens, Atsushi NitandaMarkov Decision ProcessesPolicy Gradient

  3. Monotone Neural Policy Iteration for High-Dimensional First-Order Hamilton--Jacobi--Bellman Equations

    Date pendingMinseok Kim, Yeongjong Kim, Namkyeong Cho +1Hamilton-Jacobi ReachabilityModel-Based Reinforcement Learning