We develop a statistical theory of temporal learnability in recurrent neural networks, quantifying the maximal temporal horizon
HN over which gradient-based learning can recover lag-dependent structure at finite sample size
N. The theory is built on the effective learning rate envelope
f(ℓ), a function that captures how gating mechanisms and adaptive optimizers jointly shape the coupling between state-space dynamics and parameter updates during Backpropagation Through Time. Under heavy-tailed (
α-stable) fluctuations, where empirical averages concentrate at rate
N−1/κα with
κα=α/(α−1), the interplay between envelope decay and statistical concentration yields explicit scaling laws for the growth of
HN: logarithmic, polynomial, and exponential temporal learning regimes emerge according to the decay law of
f(ℓ). These results identify envelope decay as the key determinant of temporal learnability. Slower attenuation of
f(ℓ) enlarges
HN, while heavy-tailed fluctuations compress it by weakening statistical concentration. Moreover, envelope geometry outweighs dataset size: slowing the envelope's decay enlarges
HN more than adding data, so more complex architectures that realize slower-decaying envelopes can be more data-efficient than simpler ones. Experiments across multiple gated architectures and optimizers corroborate these structural predictions.