stat.MLMay 7, 2026

Beyond the Independence Assumption: Finite-Sample Guarantees for Deep Q-Learning under ττ-Mixing

Authors: Leon Halgryn, Sophie Langer, Janusz M. Meylahn, E. Moritz Hahn

Organizations: University of Twente · Department of Applied Mathematics University of Twente · Ruhr-Universität Bochum · Faculty of Mathematics Ruhr-Universität Bochum

Abstract

Finite-sample analyses of deep Q-learning typically treat replayed data as independent, even though it is sampled from temporally dependent state-action trajectories. We study the Deep Q-networks (DQN) algorithm under explicit dependence by modelling the minibatches used for updating the network as ττ-mixing. We show that this assumption holds under certain dependence conditions on the underlying trajectories and the mechanism used to sample minibatches. Building on this observation, we extend statistical analyses of DQN with fully connected ReLU architectures to dependent data. We formulate each update as a nonparametric regression problem with ττ-mixing observations and derive finite-sample risk bounds under this dependence structure. Our results show that temporal dependence leads to a degradation in the statistical rate by inducing an additional dimensionality penalty in the rate exponent, reflecting the reduced effective sample size of ττ-mixing data. Moreover, we derive the sample complexity of DQN under tautau-mixing from these risk bounds. Finally, we empirically demonstrate on standard Gymnasium environments that the independence assumption is systematically violated and that replay sampling yields approximately exponentially decaying correlations, supporting our theoretical framework.

Explore similar work

CardsList