DeepStock: Reinforcement Learning with Policy Regularizations for Inventory Management
Authors: Yaqi Xie, Xinru Hao, Jiaxi Liu, Will Ma, Linwei Xin, Lei Cao, Yidong Zhang
Organizations: Booth School of Business, University of Chicago, Chicago, USA. · Taobao & Tmall Group, Hangzhou, China. · School of Economics, Sichuan University, Chengdu, China. · Graduate School of Business, Columbia University, New York, USA. · School of Operations Research and Information Engineering, Cornell University, Ithaca, USA.
Deep Reinforcement Learning (DRL) provides a general-purpose methodology for training inventory policies that can leverage big data and compute. However, off-the-shelf implementations of DRL have seen mixed success, often plagued by high sensitivity to the hyperparameters used during training. In this paper, we show that by imposing policy regularizations, grounded in classical inventory concepts such as "Base Stock", we can significantly accelerate hyperparameter tuning and improve the final performance of several DRL methods. We report details from a 100% deployment of DRL with policy regularizations on Alibaba's e-commerce platform, Tmall. We also include extensive synthetic experiments, which show that policy regularizations reshape the narrative on what is the best DRL method for inventory management.
Figures & tables
Figure 1 : Validation and Testing Loss Gaps for the 6 combinations of DRL Method and Policy Regularization.
Figure 2 : Validation Loss Gaps for the top-5 hyperparameter configurations, shown for each combination of DRL Method and Policy Regularization, in Setting 1.
Figure 3 : Critic loss and states during the first 10,000 timesteps in Setting 1 under a hyperparameter configuration tuned for DDPG None . The training starts after 500 transitions are collected via random actions. Each dot in the two bottom subfigures corresponds to a single visit to state s=(It,xt) .
Figure 4 : Testing and Validation Loss Gaps for DS with Base Regularization in Setting 4.
Table 1 : Stockout Rates (reported in % of days) and Turnover Times (reported in days) on the test data, comparing all policies to π\textscDDPG,\textscBoth . Positive values mean worse than π\textscDDPG,\textscBoth ; negative values mean better.
SKU Category
A+
A
B
C
D
Z
International SKU’s (change in days)
-1.36
-0.82
-1.02
-1.27
–
–
Domestic SKU’s (change in days)
-4.04
-3.81
-2.93
-2.23
-2.13
-0.64
Table 2 : Changes in average Turnover Time without changing average Stockout Rate. SKU’s are categorized by sales volume from A+ (fastest-moving) to Z (long-tail). The time period is April 2025.
Figure 5 : Evolution of average Turnover Time for international SKU’s during July–August, in 2024 and 2025. All numbers are normalized relative to the maximum average Turnover Time encountered in either year.
INDEP
AR(1)
policy
‘MlpPolicy’
learning_rate = get_linear_fn
learning_rate
tune.loguniform (3e-4, 9e-3)
tune.loguniform (6e-4, 9e-3)
lr_min
tune.loguniform (6e-5, 3e-4)
tune.loguniform (9e-5, 6e-4)
lr_fraction
0.95
batch_size
tune.choice ([128, 256, 512])
tau
tune.loguniform (1e-3, 1e-2)
Table 3: Hyperparameter values or ranges of DDPG .
INDEP
AR(1)
policy
‘MlpPolicy’
learning_rate = get_linear_fn
learning_rate
tune.loguniform (5e-5, 9e-3)
tune.loguniform (6e-4, 1e-2)
lr_min
tune.loguniform (1e-5, 5e-5)
tune.loguniform (9e-5, 6e-4)
lr_fraction
0.9
clip_range = get_linear_fn
clip_range
tune.uniform (0.1, 0.3)
lr_fraction
0.9
Table 4: Hyperparameter values or ranges of PPO .
INDEP
AR(1)
IID
policy
‘MlpPolicy’
neurons_per_hidden_layer
tune.choice([[32, 32], [64, 64], [128, 128]])
tune.choice([[16, 16], [32, 32], [64, 64]])
inner_layer_activation
‘elu’
output_layer_activation
‘relu’
batch_size
tune.choice ([5, 10])
tune.choice ([5, 10, 20])
learning_rate
tune.loguniform (4e-4, 1e-1)
Table 5: Hyperparameter values or ranges of DS .
DRL Method
DDPG
PPO
DS
Policy Regularization
None
Base
None
Base
None
Base
Validation Loss
0.9340
0.9264
0.9045
0.8917
1.2371
1.0454
Validation Loss Gap (%)
7.73
6.86
4.33
2.85
42.69
20.57
Testing Loss
0.9443
0.9356
0.9104
0.8966
1.2540
1.0959
Testing Loss Gap (%)
8.92
7.92
5.01
3.41
44.64
26.40
Table 6 : Validation and Testing Loss and corresponding Loss Gaps relative to the benchmark loss of 0.8670 in Setting 1.
DRL Method
DDPG
PPO
DS
Policy Regularization
None
Base
None
Base
None
Base
Validation Loss
2.3139
2.2647
2.4428
2.3069
2.2561
2.2550
Validation Loss Gap (%)
1.32
-0.83
6.96
1.01
-1.21
-1.26
Testing Loss
2.5735
2.4557
2.7257
2.5273
2.4455
2.4342
Testing Loss Gap (%)
12.69
7.53
19.35
10.66
7.08
6.58
Table 7 : Validation and Testing Loss and corresponding Loss Gaps relative to the benchmark loss of 2.2838 in Setting 2.
DRL Method
DDPG
PPO
DS
Policy Regularization
None
Base
None
Base
None
Base
Validation Loss
2.2341
2.3236
2.3515
2.4747
2.1643
2.2141
Validation Loss Gap (%)
-2.18
1.74
2.97
8.36
-5.23
-3.05
Testing Loss
2.6275
2.4727
2.6664
2.5319
2.5017
2.5060
Testing Loss Gap (%)
15.05
8.27
16.75
10.86
9.54
9.73
Table 8 : Validation and Testing Loss and corresponding Loss Gaps relative to the benchmark loss of 2.2838 in Setting 3.
T=5
T=17
T=33
T=65
T=129
∣Dtrain∣=5
Benchmark Loss
2.3669
2.3726
2.3792
2.3884
2.3764
Validation Loss
1.8045
1.4850
2.4117
2.6374
2.2657
Validation Loss Gap (%)
-23.76
-37.41
1.37
10.43
-4.66
Testing Loss Gap
2.8733
2.7180
2.8351
2.5135
2.5745
Testing Loss Gap (%)
21.40
14.56
19.16
5.24
8.33
∣Dtrain∣=10
Benchmark Loss
2.2569
2.2672
2.2712
2.2728
2.2640
Table 9 : Validation and Testing Loss and corresponding Loss Gaps for DS with Base regularization in Setting 4.
TUM School of Management, Technical University of Munich · Esade, Ramon Llull University · Munich Data Science Institute (MDSI), Technical University of Munich