cs.LGSep 30, 2026

Learning Under Forgetting: Statistical Support-Selective Retention in Stochastic Training Dynamics

Authors: Fujie Gao, Zuyue Zhang, Gang Sun

Abstract

Prior work has shown that neural networks exhibit implicit biases toward low-complexity structure (e.g., spectral bias), memorization dynamics, and compression-like effects during training, but a unified dynamical account of selective retention remains incomplete. We propose Repeated Reinforcement with Persistent Forgetting (RPF) dynamics, a minimal framework in which repeated exposure reinforces patterns and structures that recur in the data, while persistent forgetting attenuates learned information. This view treats forgetting not merely as a failure mode, but as a selection mechanism. We build the theory in three successive layers. First, in an independent-feature model, we derive an exposure-selective survival law and a support-dependent retention boundary characterizing which patterns persist under forgetting. Second, in a shared-parameter model, we show that forgetting induces spectral filtering over covariance modes, preserving strongly supported shared components while suppressing weak ones. Third, under small-step and norm/coding approximations, we show how RPF dynamics induce an implicit trade-off between data fitting and the cost of stored information, yielding Minimum Description Length (MDL)-like compression. Controlled experiments provide evidence for this reinforcement--forgetting selection mechanism in scalar memories and a nonlinear shared network. Joint reinforcement and attenuation interventions shift conditional retention, while matched exposure counts reveal forgetting-dependent effects of reinforcement timing and changes in the composition of the retained set. Together, these results show that repeated reinforcement and persistent forgetting jointly provide a controllable source of inductive bias beyond neural architecture and scale.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias

    May 27, 2026Mohua Das, Pierfrancesco Beneventano, Shibshankar Dey +2Neural Network TrainingBatch

  2. Not Just After One: Sleep-Inspired Replay Prevents Catastrophic Forgetting After Sequential Tasks

    Jun 7, 2026Anthony Bazhenov, Jean Erik Delanois, Giri P. KrishnanCatastrophic ForgettingContinual Learning

  3. The Rules-and-Facts Model for Simultaneous Generalization and Memorization in Neural Networks

    Mar 26, 2026Gabriele Farné, Fabrizio Boncoraglio, Lenka ZdeborováMemorizationSingular Learning Theory