Organizations: Mathematical and Algorithmic Science Laboratory, Huawei Paris Research Center, 92100 Boulogne-Billancourt, France · Laboratoire d’Informatique Gaspard Monge, Université Gustave Eiffel, 77420 Champs-sur-Marne, France
We develop an information-theoretic framework for generalization in next-token prediction under temporally dependent data. We consider independent trajectories generated by finite-memory Markov processes and distinguish algorithmic dependence, quantified by mutual information, from temporal dependence, characterized by mixing. For cross-entropy loss, we derive an expected generalization bound using the Donsker--Varadhan variational representation and a McDiarmid-type concentration inequality for Markov chains. A refinement captures the joint effect of context length and temporal mixing through the mixing properties of the history-state process. We then extend the bound through a rate--distortion formulation, replacing mutual information with the minimum information rate required to represent the learned model within a prescribed distortion in the generalization gap, yielding informative guarantees for deterministic algorithms over continuous hypothesis spaces. For margin-based prediction, we derive explicit bounds for linear and self-attention next-token predictors via noisy low-dimensional compression, revealing the roles of context length, model complexity, sample size, margin, and temporal mixing. Experiments on TinyStories and ETTh2 show that longer contexts can reduce both training and test losses, but typically reduce training loss more, enlarging the generalization gap. A complementary ETTh2 analysis identifies an effective predictive-memory scale near 24 hours, with no statistically supported improvement beyond this scale, offering a plausible explanation for test-performance saturation at larger contexts.
Figures & tables
Figure 1: Effect of context length on next-token prediction over TinyStories. A causal decoder-only Transformer is trained separately for ρ∈{2,4,8,16,32,64,128} while keeping the optimization protocol and total gradient-step budget fixed. Panel (a) reports the training and test cross-entropy losses, whereas panel (b) reports their difference, gen=Ltest−Ltrain . Shaded regions indicate one standard deviation across independent training realizations. Increasing the context length substantially improves next-token prediction while simultaneously enlarging the discrepancy between the training and test losses.
Figure 2: Effect of context length on next-token prediction over ETTh2. The context length is varied over ρ∈{2,4,8,16,32,64,128} while keeping the optimization protocol and the total budget of 15000 gradient updates fixed. Panel (a) reports the training and test cross-entropy losses, while panel (b) reports gen=Ltest−Ltrain . Curves are averaged over 5 independent random seeds, and shaded regions denote 95% confidence intervals for the corresponding means.
Figure 3: Empirical estimation of the effective predictive-memory scale of ETTh2. Panel (a) reports the coarse search over context lengths ranging from 1 to 168 hours. Panel (b) reports the refined search over 20 – 36 hours with a two-hour resolution. Both experiments use four expanding-window temporal cross-validation folds and three random seeds, with error bars representing one standard deviation. The minimum mean out-of-sample NLL is attained at ρ=24 hours in both searches. The refined comparison finds no statistically supported predictive improvement from increasing the context beyond this value, while several moderately larger contexts, including ρ=34 , remain close to the optimum. The results therefore indicate a dominant predictive-memory scale around one day together with a broader near-optimal predictive plateau, rather than a sharp empirical memory cutoff.