stat.MLSep 22, 2024

Exploiting Exogenous Structure for Sample-Efficient Reinforcement Learning

Authors: Jia Wan, Sean R. Sinclair, Devavrat Shah, Martin J. Wainwright

Organizations: Laboratory for Information and Decision Systems, Massachusetts Institute of Technology · Department of Industrial Engineering and Management Sciences, Northwestern University

Abstract

We study a structured class of Markov Decision Processes, known as Exo-MDPs, in which the state space is partitioned into exogenous and endogenous components. Exogenous states evolve stochastically, independent of the agent's actions, while endogenous states evolve deterministically based on both state components and actions. Exo-MDPs capture many operations research settings, including inventory control, resource management, and ride-sharing. Our first contribution is structural: we establish a representational equivalence between discrete MDPs, Exo-MDPs, and discrete linear mixture MDPs. Our second contribution is statistical. We characterize the minimax regret of learning in Exo-MDPs when the effective dimension r is small relative to the endogenous state and action spaces. When the exogenous states are unobserved, we prove matching upper and lower regret bounds of order Θ(HrK)Θ(Hr \sqrt{K}) over KK episodes of horizon HH, where rr is the effective dimension of the Exo-MDP. When exogenous states are observed, the minimax regret improves to Θ(HrK)Θ(H\sqrt{ r K}), revealing a Θ(r)Θ(\sqrt{r}) statistical gap due to observation of the exogenous states. These results show that Exo-MDPs decouple sample complexity from action space and endogenous state space. We validate these insights with experiments on inventory control and resource allocation.

Figures & tables

Explore similar work

CardsList
  1. Minimax Optimal Variance-Aware Regret Bounds for Multinomial Logistic MDPs

    May 19, 2026Pierre Boudart, Pierre Gaillard, Alessandro RudiMarkov Decision ProcessesMinimax

  2. Minimax PAC Bounds for Learning in Exogenous Contextual MDPs

    Jun 23, 2026Corentin Pla, Hugo Richard, Marc Abeille +1Markov Decision ProcessesMinimax