Deep Weighted Bellman Residual Minimization for Q∗ Estimation
Authors: Lican Kang, Jerry Zhijian Yang, Cheng Yuan, Chen Zhong
Organizations: Institute for Math and AI, Wuhan University, Wuhan, 430072, China · Hubei Key Laboratory of Computational Science, Wuhan University, Wuhan, 430072, China · School of Artificial Intelligence, Wuhan University, Wuhan, 430072, China · School of Mathematics and Statistics, Wuhan University, Wuhan, 430072, China
Off-policy evaluation is a foundational component of offline reinforcement learning, aiming to assess and optimize policy performance using pre-collected datasets. However, such datasets often suffer from pronounced challenges, including distribution shift, Q-value overestimation, and low sample utilization efficiency. To address these issues, this paper introduces a weighted Bellman residual minimization framework that incorporates density ratio weighting by effectively integrating expert demonstrations with behavioral data. The proposed weighting scheme departs from the conventional completeness assumption commonly imposed in the theoretical analysis of deep reinforcement learning. We establish a sharp convergence rate for density ratio estimation and derive the convergence rate for the excess risk of resulting deep Q∗ estimator. Extensive empirical evaluations demonstrate that, compared to existing methods, our method achieves significant improvements in numerical performance and policy generalization, providing specific guidance for the rational utilization of expert demonstrations.
Figures & tables
Parameters
Experiment 1
Experiment 2
Experiment 3
F
diag(0.5,0.5)
diag(0.6,0.6)
diag(0.7,0.7)
G
diag(0.5,0.5)
diag(0.4,0.4)
diag(0.3,0.3)
R
diag(0.25,0.25)
diag(0.16,0.16)
diag(0.09,0.09)
Table 1: System Parameters for Different Experiments
Algorithm
Experiment 1
Experiment 2
Experiment 3
DQN
85.09
84.32
84.57
DDPG
84.74
85.49
86.20
MABO
85.02
85.64
84.64
Ours
86.65
86.74
87.10
Table 2: Means of Linear Environment Results
Figure 1: Numerical results of Experiment 1. The theoretical optimal value is 87.84. Left: Cumulative reward variation with training epochs; Right: Evaluation reward after training completion.
Figure 2: Numerical results of Experiment 2. The theoretical optimal value is 87.84. Left: Cumulative reward variation with training epochs; Right: Evaluation reward after training completion.
Figure 3: Numerical results of Experiment 3. The theoretical optimal value is 87.84. Left: Cumulative reward variation with training epochs; Right: Evaluation reward after training completion.
Figure 4: Numerical analysis results of MountainCarContinuous Experiment. The theoretical optimal value is 100. Left: Cumulative reward variation with training epochs; Right: Evaluation reward after training completion.
Algorithm
DQN
DDPG
MABO
Ours
Rewards
94.67
95.29
94.78
95.51
Table 3: Means of MountainCarContinuous Environment Results
This paper investigates the off-policy evaluation (OPE) problem from a distributional perspective. Rather than focusing solely on the expectation of the total return, as in most existing OPE methods, we aim to estimate the entire return distribution. To this end, we introduce a quantile-based approach for OPE using deep quantile process regression, presenting a novel algorithm called Deep Quantile Process regression-based Off-Policy Evaluation (DQPOPE). We provide new theoretical insights into the deep quantile process regression technique, extending existing approaches that estimate discrete quantiles to estimate a continuous quantile function. A key contribution of our work is the rigorous sample complexity analysis for distributional OPE with deep neural networks, bridging theoretical analysis with practical algorithmic implementations. We show that DQPOPE achieves statistical advantages by estimating the full return distribution using the same sample size required to estimate a single policy value using conventional methods. Empirical studies further show that DQPOPE provides significantly more precise and robust policy value estimates than standard methods, thereby enhancing the practical applicability and effectiveness of distributional reinforcement learning approaches.
Qi Kuang, Chao Wang, Yuling Jiao +1
School of Statistics and Data Science, and Philosophy and Social Sciences Laboratory of Data Science in Finance and Economics at the Ministry of Education,2026 Jiangxi University of Finance and Economics · School of Statistics and Data Science, Shanghai University of Finance and EconomicsApr · School of Artificial Intelligence, and Hubei Key Laboratory of Computational Science, 24 Wuhan University
We present a novel theoretical framework, Q-MMR, for off-policy evaluation in finite-horizon MDPs. Q-MMR learns a set of scalar weights, one for each data point, such that the reweighted rewards approximate the expected return under the target policy. The weights are learned inductively in a top-down manner via a moment matching objective against a value-function discriminator class. Notably, and perhaps surprisingly, a data-dependent finite-sample guarantee for general function approximation can be established under only the realizability of Qπ, with a dimension-free bound -- that is, the error does not depend on the statistical complexity of the function class. We also establish connections to several existing methods, such as importance sampling and linear FQE. Further theoretical analyses shed new light on the nature of coverage, a concept of fundamental importance to offline RL.
Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate these ratios by enforcing occupancy-balance moments over a critic class. We propose fitted occupancy-ratio evaluation (FORE), a fitted fixed-point method that characterizes the discounted occupancy ratio through an adjoint Bellman recursion. At each iteration, FORE solves a single-level density-ratio objective on one-step-transition data, thereby projecting the adjoint Bellman image onto a log-ratio class in Kullback--Leibler (KL) divergence. Unlike analyses of fitted Q-evaluation, which typically require value-function realizability together with Bellman completeness or projected-operator stability, our central approximation condition is just realizability of the discounted occupancy ratio itself. Under this condition, the population KL-projected recursion contracts in relative entropy toward the true ratio by virtue of the adjoint Bellman operator being a KL-contraction. For the empirical recursion, we establish finite-sample regret bounds that yield convergence in KL up to log-ratio approximation error and a statistical error governed by the complexity of the ratio hypothesis class. The fitted ratio supports direct value estimation by reward reweighting, occupancy-weighted fitted Q-evaluation, and doubly robust estimation that combines the fitted ratio with a fitted Q-function. Together, these results identify discounted occupancy-ratio realizability as a sufficient condition for offline policy evaluation without any completeness assumptions.
Lars van der Laan, Nathan Kallus
Stanford University · Netflix · Cornell University