Authors: Yaniv Oren, Viliam Vadocz, Wiktor Zabka, Thomas Evers, Jan Robine, Wendelin Böhmer, Matthijs T. J. Spaan, Martha White, +2 more
Organizations: Department of Intelligent Systems, TU Delft · Centrum Wiskunde & Informatica, Amsterdam · Department of Computer Science, ETH Zürich · Department of Computing Science, University of Alberta · Information Systems, TU Eindhoven · Delft Institute of Applied Mathematics, TU Delft
Practical Reinforcement Learning (RL) algorithms learn to solve Markov Decision Processes (MDPs) through iterative policy improvement in the presence of approximate evaluation. We study policy improvement from first principles, defining optimal policy improvement as producing the best policy attainable in a single update under specified constraints. We show that optimal improvement restricted to a set of states is equivalent to solving an induced MDP, characterizing planning with an explicit or implicit model as a path towards optimal policy improvement. Because practical methods commonly solve such induced problems through iterative improvement in the form of greedification, we take steps towards optimal greedification under the central practical constraint of approximate evaluation. We formulate greedification under this constraint as probabilistic decision-making under uncertainty and derive a novel operator that is optimal with respect to the resulting objective. Empirically, the operator and its practical gradient-based approximations improve aggregate performance across GumbelAlphaZero, SAC, ReBRAC and Generalized Policy Iteration, in experiments spanning discrete and continuous actions, model-based and model-free, online and offline RL.
Figures & tables
Figure 1: Noised policy iteration: normalised Vπ(s0) at iteration 40 over 20 grid environments, 95% Gaussian CI. Model-based, discrete: Bayes Elo at 34 M training frames, mean, min. and max. over 5 training seeds in 9x9 Go. Model-free, continuous: mean normalised returns at 1 M steps on 10 DeepMind Control Suite tasks, mean and 95% Gaussian CI. Offline: normalised return at 1 M steps on 4 D4RL medium-replay tasks, mean and 95% Gaussian CI.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Policy Iteration experiments with different policy improvement operators and uncertain evaluation, mean and 95% Gaussian CI. Left: Example learning curves. Center: Sample-efficiency vs. starting noise magnitude, measured as Area Under the Curve (AUC) of learning curves until 100 iterations (higher is better). Right: AUC vs. hyperparameter combination. The performance landscape exhibits a single broad ridge of high performance, with performance degrading smoothly away from the ridge.
Figure 3: Policy Iteration with a 1-step Bellman update ( k=1 ), otherwise identical to Figure 2 .
Figure 4: Playing strength in training on 9×9 Go with M=64 . Every checkpoint of every agent and every seed plays every other resulting in a single set of Elo scores. Bold curves are the mean over 5 training seeds per agent. Faint curves are the seeds themselves, with 95% BayesElo CI.
Figure 5: Playing strength per agent with the same set of DNNs. Bayes elo and 95% CIs.
Figure 6: Left: Variance based on one-step TD residual proxies for evaluation error on different opening-book positions under one network on 9×9 Go with 95% percentile bootstrap over positions. Right: the shape of the Qϕ residual, on a log density scale. In grey the raw pooled residual, a scale mixture over actions of differing σ . In blue is the same data centered and scaled per position; dashed is the fitted N(0,1) .
Figure 7: Per-environment learning curves on the 10 DeepMind Control Suite tasks, showing unnormalised episodic return against environment steps. Mean across 10 seeds and 95% interval.
ReBRAC +Iopt
Dataset
per-dataset τ2,β
single pair τ2=2 , β=0.5
ReBRAC
halfcheetah-medium-replay
50.86±1.09
50.30±1.43
49.65±0.28
hopper-medium-replay
87.00±8.80
87.00±8.80
88.24±6.80
walker2d-medium-replay
84.69±3.19
84.69±3.19
78.75±3.91
ant-medium-replay
40.17±7.89
40.17±7.89
17.09±3.05
Appendix
Table 1: Per-environment offline RL results on the four D4RL medium-replay datasets. Entries are the D4RL score of the final policy evaluated over 56 episodes, mean and 95% Gaussian CI over 20 runs per arm per environment ( 10 for the single-pair setting on halfcheetah). Bold marks the highest mean in each row together with every entry that is not significantly worse than it, at p<0.05 under a two-sided Welch t -test over runs ( Welch, 1947 ) ; several bold entries in a row therefore indicate a tie.
Experiment
Number of Seeds per Agent
Policy iteration
100 seeds ×20 environments
Model-based, discrete
5 training seeds
Model-free, continuous
10 seeds ×10 environments
Offline
20 seeds ×4 environments
Appendix
Table 2: Number of seeds included in each experiment.
Component
Parameter
Value
Iopt
Advantage prior variance ( τ2 )
(0.17)2
KL regularization strength ( β )
0.07
Ireg
Temperature ( η )
4.89
Environments
Grid size ( ∣S∣×∣A∣ )
25×5
Noise schedule
Initial noise bound ( a0 )
50
(Eq. 135 )
Decay ( γdec )
0.998
Appendix
Table 3: Hyperparameters for the Policy Iteration experiments.
Component
Parameter
Value
Iopt
Advantage prior variance ( τ2 )
0.08
KL regularization strength ( β )
0.01
Ireg (GumbelMCTS)
cvisit
50
cscale
1.0
PUCT (AlphaZero)
cinit
1.25
cbase
19652
Appendix
Table 4: Hyperparameters for the discrete-action model-based GAZ experiments.
Component
Parameter
Value
Iopt
Advantage prior variance ( τ2 )
0.4
Num. actions sampled for Vϕi(s)
32
SAC
Optimizer
Adam
Actor learning rate
3⋅10−4
Critic and entropy learning rate
10−3
Batch size
256
Appendix
Table 5: Hyperparameters for the online classical continuous control SAC experiments
Center for Applied Computing, Faculty of Information Technology and Electrical Engineering, University of Oulu, Finland · Dept. of Advanced Computing Sciences, Maastricht University, the Netherlands