Deep Weighted Bellman Residual Minimization for Q∗ Estimation
Authors: Lican Kang, Jerry Zhijian Yang, Cheng Yuan, Chen Zhong
Organizations: Institute for Math and AI, Wuhan University, Wuhan, 430072, China · Hubei Key Laboratory of Computational Science, Wuhan University, Wuhan, 430072, China · School of Artificial Intelligence, Wuhan University, Wuhan, 430072, China · School of Mathematics and Statistics, Wuhan University, Wuhan, 430072, China
Off-policy evaluation is a foundational component of offline reinforcement learning, aiming to assess and optimize policy performance using pre-collected datasets. However, such datasets often suffer from pronounced challenges, including distribution shift, Q-value overestimation, and low sample utilization efficiency. To address these issues, this paper introduces a weighted Bellman residual minimization framework that incorporates density ratio weighting by effectively integrating expert demonstrations with behavioral data. The proposed weighting scheme departs from the conventional completeness assumption commonly imposed in the theoretical analysis of deep reinforcement learning. We establish a sharp convergence rate for density ratio estimation and derive the convergence rate for the excess risk of resulting deep Q∗ estimator. Extensive empirical evaluations demonstrate that, compared to existing methods, our method achieves significant improvements in numerical performance and policy generalization, providing specific guidance for the rational utilization of expert demonstrations.
Figures & tables
Parameters
Experiment 1
Experiment 2
Experiment 3
F
diag(0.5,0.5)
diag(0.6,0.6)
diag(0.7,0.7)
G
diag(0.5,0.5)
diag(0.4,0.4)
diag(0.3,0.3)
R
diag(0.25,0.25)
diag(0.16,0.16)
diag(0.09,0.09)
Table 1: System Parameters for Different Experiments
Algorithm
Experiment 1
Experiment 2
Experiment 3
DQN
85.09
84.32
84.57
DDPG
84.74
85.49
86.20
MABO
85.02
85.64
84.64
Ours
86.65
86.74
87.10
Table 2: Means of Linear Environment Results
Figure 1: Numerical results of Experiment 1. The theoretical optimal value is 87.84. Left: Cumulative reward variation with training epochs; Right: Evaluation reward after training completion.
Figure 2: Numerical results of Experiment 2. The theoretical optimal value is 87.84. Left: Cumulative reward variation with training epochs; Right: Evaluation reward after training completion.
Figure 3: Numerical results of Experiment 3. The theoretical optimal value is 87.84. Left: Cumulative reward variation with training epochs; Right: Evaluation reward after training completion.
Figure 4: Numerical analysis results of MountainCarContinuous Experiment. The theoretical optimal value is 100. Left: Cumulative reward variation with training epochs; Right: Evaluation reward after training completion.
Algorithm
DQN
DDPG
MABO
Ours
Rewards
94.67
95.29
94.78
95.51
Table 3: Means of MountainCarContinuous Environment Results
School of Statistics and Data Science, and Philosophy and Social Sciences Laboratory of Data Science in Finance and Economics at the Ministry of Education,2026 Jiangxi University of Finance and Economics · School of Statistics and Data Science, Shanghai University of Finance and EconomicsApr · School of Artificial Intelligence, and Hubei Key Laboratory of Computational Science, 24 Wuhan University