cs.LGSep 30, 2026

Learning Infinite-Horizon Average-Reward CMDPs via State Augmentation

Authors: Kihyun Yu, Seoungbin Bae, Dabeen Lee

Organizations: KAIST · Seoul National University

Abstract

We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the weakly communicating assumption. Existing high-probability guarantees for this setting either require computationally inefficient algorithms or have suboptimal dependence on the number of interactions TT. We propose, to the best of our knowledge, the first computationally efficient algorithm that achieves O~(T)\widetilde{\mathcal{O}}(\sqrt{T}) regret and cumulative constraint violation with high probability in the tabular setting. The T\sqrt{T} dependence is optimal up to logarithmic factors. Our approach incorporates cumulative constraint violation into the state and defines a reshaped reward through differences of a Huber potential. The added state determines the penalty on further violations while the reward function remains fixed on the augmented state space. Since the added state has known deterministic dynamics, only the original transition kernel needs to be estimated. The bounded slope of the Huber potential keeps the per-step reward bounded, and the potential differences telescope to relate the reshaped return to the original cumulative reward and the terminal potential. These properties allow us to apply finite-horizon approximation and optimistic value iteration with clipping, as used in unconstrained average-reward MDPs, without worsening the regret rate in TT.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Slater's Condition in Online CMDPs with Stochastic and Adversarial Constraints

    Sep 24, 2025Francesco Emanuele Stradi, Eleonora Fidelia Chiefari, Matteo Castiglioni +2Markov Decision ProcessesConstrained RL