cs.LGMar 20, 2026

DeepStock: Reinforcement Learning with Policy Regularizations for Inventory Management

Authors: Yaqi Xie, Xinru Hao, Jiaxi Liu, Will Ma, Linwei Xin, Lei Cao, Yidong Zhang

Organizations: Booth School of Business, University of Chicago, Chicago, USA. · Taobao & Tmall Group, Hangzhou, China. · School of Economics, Sichuan University, Chengdu, China. · Graduate School of Business, Columbia University, New York, USA. · School of Operations Research and Information Engineering, Cornell University, Ithaca, USA.

Abstract

Deep Reinforcement Learning (DRL) provides a general-purpose methodology for training inventory policies that can leverage big data and compute. However, off-the-shelf implementations of DRL have seen mixed success, often plagued by high sensitivity to the hyperparameters used during training. In this paper, we show that by imposing policy regularizations, grounded in classical inventory concepts such as "Base Stock", we can significantly accelerate hyperparameter tuning and improve the final performance of several DRL methods. We report details from a 100% deployment of DRL with policy regularizations on Alibaba's e-commerce platform, Tmall. We also include extensive synthetic experiments, which show that policy regularizations reshape the narrative on what is the best DRL method for inventory management.

Figures & tables

Explore similar work

CardsList
  1. InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees

    May 1, 2026Chenyu Huang, Jianghao Lin, Zhengyang Tang +4Inventory ControlLarge Language Model Policy Optimization

  2. Hard Constraints, Smooth Gradients: Learning Feasible Inventory Policies via Differentiable Projection

    Aug 3, 2026Patrick Helm, Jan-Niklas Doerr, Joren Gijsbrechts +1Mixed-Integer ProgrammingInequality Constraints

  3. Automated Design of Inventory Policy with Large Language Models: An Exploratory Study

    Sep 8, 2026Fenghua Yang, Preet Baxi, Yi Zhang +4Inventory ControlLarge Language Model Policy Optimization