cs.LGJun 23, 2026

Bias-Controlled Primal-Dual Natural Actor-Critic: Optimal Rates for Constrained Multi-Objective Average-Reward RL

Authors: Ankur NaskarSwetha GaneshVaneet Aggarwal

Organizations: Indian Institute of Science, Bengaluru, India · Purdue University, USA

Abstract

Many reinforcement learning (RL) problems in the infinite-horizon average-reward setting require optimizing multiple conflicting objectives while satisfying multiple safety constraints. A common approach is concave scalarization, where the agent maximizes a utility f(Jr1π,,JrMπ)f(J^π_{r_1}, \ldots, J^π_{r_M}) subject to a scalarized constraint g(Jc1π,,JcNπ)0g(J^π_{c_1}, \ldots, J^π_{c_N}) \ge 0, where JrmπJ^π_{r_m} and JcnπJ^π_{c_n} denote the average-reward and cost under policy ππ. However, the nonlinearity of ff and gg introduces bias in policy-gradient and actor-critic methods, since gradients must be evaluated using noisy estimates of Jπ,J^π, and E[f(Jπ)]f(E[Jπ]), \mathbb{E}[\partial f(J^π)] \neq \partial f(\mathbb{E}[J^π]), and this bias propagates through both primal and dual updates. We propose an MLMC-based primal-dual Natural Actor-Critic algorithm for average-reward MDPs that controls bias in scalarized objectives, constraint evaluation, and actor-critic estimation without requiring mixing-time knowledge. We show that the algorithm achieves optimal global convergence and constraint-violation rates of O~(1/T)\tilde{O}(1/\sqrt{T}). To our knowledge, this is the first result establishing optimal convergence for concave scalarized multi-objective RL in the average-reward setting, both with and without constraints, and the first to do so without mixing-time information even in the absence of scalarization.

Explore similar work

CardsList