Neural Network Training Dynamics

Latest papers 349

Sep 28, 2026cs.LG

On Parameter Symmetries and Conservation Laws in Gradient Flow

Parameter space symmetries and conservation laws play an important role in understanding the loss landscapes and implicit biases of neural networks. Inspired by Noether's theorem in physics, prior works have sought to derive conservation laws under gradient flow from parameter symmetries, but the scope and limitations of this connection remain unclear. We develop a unified geometric framework that clarifies the precise relationship between the two notions, including the conditions under which symmetries correspond to conservation laws. We introduce a notion of compositional identifiability and use it to establish a general inheritance principle for complete characterizations of symmetries and conservation laws in multilayer networks. We apply the framework to multi-head and grouped-query attention, polynomial neural networks, and square deep linear networks.
Sep 28, 2026cs.LG

LLN: Learnable Lens Networks for Parameter-Efficient Long-Horizon Dynamical Prediction

Explicit residual connections of the form (x+f(x)), often combined with normalization layers, have become a standard strategy for training very deep neural networks. However, residual addition primarily provides an algebraic shortcut for gradient propagation, while leaving the evolution of feature geometry across layers largely unconstrained. We introduce Learnable Lens Networks (LLN), a physics-inspired architecture that replaces direct feature-space residual accumulation with learnable optical transport in an augmented position-angle phase space. Each layer alternates between free propagation, which provides an implicit transport path, and a learnable lens field that performs nonlinear trajectory transformation and focusing. Theoretically, we establish that LLN transport is globally invertible and volume-preserving for any differentiable lens field, with the implemented coordinate-wise Gaussian transport further satisfying symplecticity. Importantly, these structural constraints do not limit expressivity: with unrestricted embeddings and readouts, LLN retain universal approximation of continuous end-to-end maps. Experiments across diverse dynamical systems demonstrate that LLN improves long-horizon prediction while using substantially fewer parameters than same-depth comparators. Further analysis reveals stable depth-wise gradient transport and interpretable learned dynamics under the coupled propagation and refraction design.
Sep 27, 2026cs.LG

Scalable Attribution and Control of Model Behavior During Training

Attributing and controlling model behavior during training requires identifying each example's contribution quickly enough to act before the next update. However, examples in the same training batch can produce similar behavioral changes, making their individual contributions difficult to distinguish. We address this ambiguity through mutual information, accounting for interference within the batch by quantifying how much the combined behavioral change reveals about each example's contribution. We show that this mutual information is a logarithmic function of Behavioral Gradient Uniqueness (BGU). BGU gives the information measure its geometric interpretation. Our Batch-Space Ghost (BS-Ghost) algorithm makes these scores practical inside the training loop through shared computation in batch space, without storing model-sized example gradients. On a complete 1,000-example Qwen2.5-7B-Instruct workload, our BS-Ghost implementation adds 27 seconds (8.0%) to 5.5 minutes of ordinary training. Removal and retraining demonstrate that BGU identifies data that causally shapes final behavior. At each training step, signed information identifies which examples strengthen or weaken the target behavior, explaining how behavior develops during training. Signed information also enables cheap intervention during training: it predicts how changing example weights will affect behavior in the next update. We then use these predictions to choose weights that steer behavior toward a desired target. This makes our framework a practical foundation for scalable oversight and verification of training pipelines and processes, helping evaluators assess model alignment, understand how it develops during training, and guide interventions that shape ongoing learning.
Sep 27, 2026cs.LG

Learning Dynamics of Continual Learning: A Unified View of Data Attribution, Forgetting, and Plasticity Loss

Modern language models are likely to be updated throughout their lifetime rather than trained once and frozen. Each update therefore participates in a recurring cycle: decide which experience to learn from, understand what that update changes, and remain capable of learning from what comes next. We show that these challenges are governed by the same evolving update--behavior interaction. We derive a token- and layer-wise decomposition of how learning from one token changes another prediction. By separating the softmax force, shared readout geometry, and residual connections, it exposes two interaction channels and yields a forward-computable approximation. Following this interaction through time reveals a unified picture of continual adaptation. Positive interaction identifies useful experience; negative interaction produces either concentrated collision or accumulated erosion; over longer horizons, updates reshape the shared geometry mediating future learning signals, reducing their transmission. These predictions lead to effective data selection, mechanism-specific controls for interference, and a readout-based diagnostic of future learnability whose degradation predicts the benefit of restoring the readout. Across models and training regimes, the same local interaction thus explains both what an update changes now and how learning today changes what can be learned tomorrow. This view connects data attribution, forgetting, and plasticity loss as distinct regimes of the same evolving learning dynamics.
Sep 24, 2026cs.LG

Tracking States or Tracking Cosets? An Algebraic Account of Learned State Tracking

State tracking requires composing a sequence of updates, but accuracy alone does not reveal what a model has learned. We study neural networks trained to predict the running product of group elements. We identify quotient solutions in Transformers, where models recover the quotient class while predicting nearly uniformly among its members. The reciprocal of class size predicts partial accuracy without a fitted parameter, extending parity-based accounts to non-parity quotients. Our baseline Transformers' predictions change little under prefix reordering beyond the exact-tracking frontier. We prove that, for finite groups under uniform i.i.d. full-group inputs, optimal order-blind exact accuracy converges to the reciprocal of abelianization class size as prefix length grows, consistent with the observed abelianization plateaus. Sequential updates permit more: any partition into right cosets of a subgroup, normal or not, survives sequential updates. In our census of standard Transformers, every recovered coset partition comes from a normal subgroup, whereas parameter-matched recurrent networks pass through both normal and non-normal right-coset stages during training. On A5A_5, we identify low-dimensional subspaces of the recurrent state that encode non-normal cosets. In the three-dimensional cases, coset mean vectors form approximate dodecahedra, and swapping the state components in these subspaces transfers the donor's coset state through a shared input suffix. Our results connect partial accuracy, learning stages, and internal computation through the subgroup cosets that models learn to track.
Sep 24, 2026cs.LG

An Analytical Theory of Auxiliary Learning

Auxiliary learning is an optimization paradigm in which a neural network's performance on a target task is improved by jointly training it on additional tasks. However, the mechanisms behind this improvement remain poorly understood. We study this problem using a teacher-student framework and derive a closed system of differential equations describing the dynamics of online stochastic gradient descent in the large-input limit. For linear networks, we obtain a closed-form expression for the generalization error to leading order in the learning rate, quantifying how task correlations and label noise determine the benefit of auxiliary learning. For non-linear activation functions, we develop a fluctuation-dissipation analytical theory that establishes a general relation linking the main and auxiliary errors to the corresponding single-task error. Numerical experiments support the theoretical predictions and show how auxiliary tasks improve generalization by balancing the forcing dynamics towards the optimal solution with gradient noise.
Sep 22, 2026cs.LG

The Dynamics of Quasiregular Neural Learning

Many learning problems combine a dominant regularity with systematic exceptions. Motivated by U-shaped learning in language acquisition, we study this interaction in controlled quasiregular regression problems where regular and exceptional solutions are explicitly known. Neural networks can partially acquire exceptions, subsequently regress toward the dominant regularity, and finally recover. This overregularization becomes substantially stronger when exceptions are rare, despite their early acquisition, but does not emerge equally across all regularities considered. Our results isolate a simple form of competition between regularities and exceptions during neural learning.
Sep 22, 2026cs.LG

An Exploratory Replica-Overlap Probe of the Grokking Transition

We trained 64 independently seeded networks in four configurations, continuing each to sustained convergence or a 40,000-epoch ceiling. We then asked whether an RSB-inspired distribution of pairwise weight overlaps changes across the grokking transition. It is the alignment step, not the overlap statistic, that determines what this registered probe can report. The registered implementation permutes hidden units without the corresponding bias and head-internal permutations and therefore does not preserve the network function. Every q_wt value computed through this alignment inherits the defect; q_fn does not, because it is computed from predictions of the unpermuted models. The numerical-precision requirement also failed, and an audit found protocol deviations. Consequently, the pre-registered rule gives no verdict: registered outcome UNDETERMINED (reason code C0_INSTRUMENT_INVALID). These data provide neither a confirmatory null nor a validated reading of the Parisi order parameter. Only frac40 cleared the 12/16 checkpoint-completeness requirement. For this configuration, a post-hoc criterion applied to the same data gave a Hartigan-dip interval containing zero (95% CI for Delta dip = [-0.017, 0.034]), whereas the overlap standard deviation increased by a factor of about 5.6. A post-hoc calibration assigns the dip test zero power at the simulated separations; the interval is therefore uninformative, not evidence of no change. The standard-deviation ratio is the only statistic here with power at the observed effect. Ensemble loss was near-flat only under the pre-specified 1% threshold. Finally, grokking rates of 0/16, 11/16 and 16/16 remain descriptive because train fraction is confounded with split identity.
Sep 17, 2026cs.NE

Position Paper: Neurotransmitters as a Missing Dimension in Artificial Neural Networks

Artificial neural networks (ANNs), as core components of modern deep learning (DL) systems, lack the adaptive flexibility and long-term stability exhibited by biological systems. This limitation largely stems from the fact that conventional ANNs rely on uniform, local, and gradient-based parameter updates, while neglecting internal learning principles that are biological mechanisms such as neurotransmitters signalling or neuroplasticity. Consequently, many existing approaches focus on architectural expansion or mathematical fine-tuning techniques such as regularisation or parameter isolation. Inspired by the superior adaptability and plasticity of mammalian brains, we posit that neuromodulation with neurotransmitters constitutes a third axis of learning, complementary to neural activity and synaptic plasticity, and should be explicitly modelled in artificial neural networks. In this positional paper, we argue that incorporating neuromodulatory principles into ANN design represents a promising and underexplored research direction, and we advocate for greater attention to this perspective in the development of adaptive and continual learning systems.
Sep 16, 2026cs.LG

Learning-Induced Dynamical Transition in Recurrent Neural Networks

Learning in recurrent neural networks can fundamentally reshape their underlying dynamics, transforming initially chaotic activity into stable task-dependent behavior. We develop a non-equilibrium dynamical mean-field theory(DMFT) to describe this transition during learning. We show that a slow feedback-driven learning process generates an evolving effective feedback strength that drives the network through a transition from chaotic to stable dynamics defined by a bifurcation of the DMFT solution. By deriving the two-time correlation function throughout learning, we identify a critical feedback strength and a corresponding learning rate dependent critical time separating these regimes. The transition arises from the progressive deformation of an effective dynamical landscape by the growing learned feedback structure. Starting from the untrained state, the theory predicts the time evolution of the network output during training and shows quantitative agreement with numerical simulations.
Sep 16, 2026cs.LG

Beyond Quadratic Loss: The Stability Phase Diagram of Adam

Loss spikes are recurrent instabilities in neural-network training and can arise from multiple mechanisms. For Adam in particular, macroscopic loss spikes have been linked to optimizer dynamics, yet how its two momentum timescales govern them remains unclear. We investigate this dependence by mapping training dynamics across the (β1,β2)(β_1,β_2) plane. Across a range of model--task settings, an approximately linear boundary, 1−β2=C(1−β1)1-β_2=C(1-β_1), separates spiky from non-spiky dynamics, whereas a one-dimensional quadratic loss produces approximately cubic slope. A one-dimensional superquadratic loss L(x)∝∣x∣nL(x)\propto|x|^n recovers the near-linear scaling and links the boundary coefficient to the effective loss exponent nn. We further show that confident cross-entropy losses develop a core--wall landscape comprising a narrow quadratic core followed by a steep wall, which produces effective superquadratic behavior at the scale of an optimizer update. Together, these results connect Adam loss spikes to both the mismatch between momentum timescales and finite-scale superquadratic loss geometry beyond the Hessian.
Sep 15, 2026cs.LG

Information Geometric Self-Organization at the Edge of Stability in High-Capacity Kernel Associative Memories

High-capacity associative memories based on Kernel Logistic Regression (KLR) exhibit exceptional storage capabilities and robustness. Previous empirical studies identified a hyperparameter regime, the "Ridge of Optimization," where attractor stability is maximized. However, the geometric nature of this regime and the optimization dynamics required to reach it have remained unclear. In this paper, we investigate the static geometry of the parameter space and the learning trajectory of Gradient Descent (GD) in KLR-trained Hopfield networks. Using the eigenvalue spectrum of the Hessian, we reveal that the Ridge corresponds to a phase boundary located adjacent to a rank-1 spectral collapse, acting as a geometric singularity where the principal curvature is massively amplified. Furthermore, we demonstrate that the learning dynamics exhibit a transient self-stabilizing behavior driven by the Edge of Stability (EoS) phenomenon. Rather than seeking flat regions, the network parameters are driven toward a state where the local curvature dynamically equilibrates near the stability limit dictated by the learning rate, allowing the optimization to survive the initial instability. We provide analytical derivations for both the rank-1 asymptotic collapse and the dynamic feedback loop governing this equilibration. These findings suggest that optimal, high-capacity memory representations are not formed in flat minima, but are dynamically sculpted at the highly curved boundaries of geometric singularities.
Sep 14, 2026cs.LG

Dimension-Corrected Hitting Times for Heavy-Tailed Spectral Emergence in Neural Optimizer Dynamics

Heavy-tailed empirical spectral densities of neural-network weight matrices are widely used as diagnostics of implicit self-regularization, but the step complexity of heavy-tail emergence remains poorly understood. We formulate spectral heavy-tail formation as a right-censored hitting-time problem: a run that does not reach a heavy-tail diagnostic within the observation horizon is treated as censored rather than discarded. In controlled full-batch teacher--student dynamics, we find that the first-step spike--bulk gap alone does not explain onset time. Instead, finite-onset regression supports a dimension-corrected spectral-gap law, (\tau_{\mathrm{HT}}\approx C\Delta_1^{-\gamma}d^\rho), with (R^2=0.683), (\gamma=0.626), and (\rho=0.772) across 330 completed runs. Right-censored lognormal accelerated-failure-time models further favor the dimension-corrected model over a gap-only model, improving AIC from 706.62 to 628.70. Theoretically, we prove that exact early loss dynamics in linear networks do not determine factor spectral tails, that Adam recurrences alone do not imply spectral redistribution, and that projected singular-basis spreading implies contraction of a spectral-tail potential and hence a dimension-corrected hitting-time bound. Empirically, projected-kernel profiles support the sufficient spreading mechanism, Adam and AdamW agree under tested grids, GD and signGD do not reach onset in the same regimes, and real pretrained Qwen2.5-0.5B and Pythia-70M transformer weights show non-Gaussian spectral-tail structure relative to matched Gaussian nulls. The result is a reproducible spectral hitting-time law with rigorous conditional theory, not a claim that Adam necessarily generates heavy tails from first principles.
Sep 14, 2026cs.LG

A Full Adam Theorem for Spectral Heavy-Tail Onset

We prove a full Adam theorem for spectral heavy-tail onset in a closed Gaussian Stein-Hermite teacher-student state-evolution model. The theorem begins with the actual full-batch Adam recurrences, derives the population gradient by Stein-Hermite calculus, proves finite-width covariance concentration, converts multi-step Adam momentum into an exact non-centered Gaussian sign kernel, controls the diagonal Adam denominator by a basis-homogenization theorem, derives a regularly varying projected update response from a Hermite edge-transfer theorem, pushes the response through the exact Gram update, and proves approximate-target KL contraction with matching upper and lower hitting bounds. The final law is (\tau_\varepsilon=\Theta(\Delta_1^{-\gamma}d^\rho\log(\Psi_0/\varepsilon))), where (\Delta_1) is the first spike-bulk spectral gap. The result is full in the following precise sense: every step from Adam's momentum and denominator to the spectral hitting law is formalized inside the closed state-evolution model. We also prove that a stronger arbitrary-gradient Adam theorem is impossible, and that exact two-step linear-network loss dynamics do not identify factor spectra or heavy-tail hitting times.
Sep 12, 2026cs.AI

Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking

Neural networks trained past memorization frequently undergo a delayed transition to generalization, a phenomenon known as grokking. Despite theoretical progress on \emph{why} this transition occurs, the quantitative structure of \emph{when} it occurs in hyperparameter space remains uncharacterized. We map the memorization-to-generalization boundary across 384 configurations of two-hidden-layer MLPs on modular arithmetic, fitting a power-law scaling relation for generalization onset time: Tgrok∝H−0.27 D−2.04 η−0.50 λ−0.64T_{\mathrm{grok}} \propto H^{-0.27}\, D^{-2.04}\, \eta^{-0.50}\, \lambda^{-0.64} (R2=0.732R^2 = 0.732; 0.8210.821 with interactions). The exponent hierarchy reveals that data complexity (D−2.04D^{-2.04}) is the dominant driver of regime transition, not model capacity (H−0.27H^{-0.27}): doubling data accelerates generalization by ∼4×{\sim}4\times, while doubling width yields only ∼1.2×{\sim}1.2\times. A sharp phase boundary at weight decay λ≳1.0\lambda \gtrsim 1.0 separates grokking from non-grokking configurations, and weight norm trajectories show monotonic compression during the transition, consistent with implicit regularization selecting low-complexity solutions. These results provide a quantitative foundation for predicting and controlling regime transitions in overparameterized networks.
Sep 11, 2026cs.LG

Learning Orthogonal Multi-Index Models Beyond Small Initialization: Incremental Learning, Competitive Dynamics and Symmetry

Recent work has identified incremental learning in shallow networks trained on single-index and multi-index models. However, existing analyses often rely on simplifying settings, such as small initialization, correlation loss, or layer-wise training. These choices reduce neuron interactions and leave some feature learning dynamics under standard initialization unexplored. We study training dynamics for polynomial-width two-layer networks learning orthogonal multi-index targets under standard initialization using polynomially many samples. We first prove that incremental learning still occurs: the loss decreases sequentially according to the Hermite expansion of the target, with lower-order components learned before higher-order components recover the individual target directions. In this standard initialization regime, training also shows a competitive reallocation of parameter mass: after the total mass fits the target mean and stabilizes, mass shifts into the target subspace and then concentrates on aligned neurons. Our theoretical analysis uses slightly modified gradient flow, while vanilla gradient descent empirically exhibits the same qualitative dynamics. Technically, we introduce a symmetry-based finite-width approximation via symmetrized networks, rather than comparing directly with an infinite-width limit. This yields better control of approximation errors and may be of independent interest.
Sep 10, 2026cs.LG

Teacher Geometry Shapes Learnability in Teacher-Student Networks

Teacher-student systems, in which a teacher neural network generates training labels so that a student neural network can learn to implement the same function, are widely used as an abstract setting to study learning. However, the structure of the teachers is often overlooked by assuming randomly-generated, normally-distributed parameters. This hides substantial variation in how learnable different teachers are. We formalize learnability as the success rate of converging to the global minimum, as a function of overparameterization, learning algorithm, student initialization distribution, and teacher geometry. We both identify an easy distribution that maximizes node dissimilarity and a hard distribution that minimizes it, and show that these two distributions induce markedly different success rates across a large range of settings and for different activation functions. To explain the gap, we study the loss landscape of small neural networks that contain two distinct kinds of suboptimal local minima, out-of-bounds (OOB) minima at the edge of the data distribution and interior minima within. Assuming infinite data and a fast readout layer, we analytically reduce the loss landscape of small networks to two dimensions, showing that the region of attraction of interior minima changes as a function of teacher structure. In larger networks, maximally dissimilar teachers induce more interior minima, while minimally dissimilar teachers induce more OOB minima. Motivated by these analyses, we show that differentially increasing the learning rate of the readout layer and decreasing the learning rate of the inner biases increases success rates. These findings provide an important step in narrowing the gap between the study of teacher-student networks and more structured functions that arise in practice.
Sep 9, 2026cs.LG

Training Trajectories Determine Circuit Removability in Annealable Soft-Prior Transformers

Soft positional priors can help small Transformers learn retrieval circuits, but it is unclear whether the resulting circuits remain functional once the prior is removed. We test this with an annealable soft-prior Transformer whose attention biases can be learned, faded, or zeroed during training and evaluation. On associative recall, unforced models perform well with the prior active (0.772±0.0200.772 \pm 0.020) but collapse at zero gate (0.095±0.0090.095 \pm 0.009). Smooth fade-to-zero training preserves high zero-gate accuracy (0.734±0.0280.734 \pm 0.028), whereas forced-zero training, hard switching, and post hoc continuation fail to recover the same effect. The pattern also appears on Markov induction. Linear regression ICL provides a boundary case because zero-gate training can learn that task directly. Mechanistic traces show that circuit consolidation occurs after the gate reaches zero, even though the responsible heads vary across seeds. These results suggest that circuit removability in small discrete retrieval tasks depends on the training trajectory, not just the final architecture.
Sep 9, 2026cs.AI

A Function-Space Approach to the Statistical Mechanics of Learning Dynamics

In the kernel regime, neural-network learning inherits its preferences from a frozen spectrum. During feature learning, this spectrum evolves, yet networks retain systematic biases toward simple, smooth directions. We develop a function-space statistical framework explaining the origin of these preferences, treating functions and their learning operators as macroscopic variables, with parameterization entering through the multiplicity of parameter configurations realizing each function. For mean-squared loss, error relaxes exactly under the evolving learning operator M=JJ∗M=JJ^\ast. Training stochasticity induces a Gaussian weight over function-space states, while parameter multiplicity contributes an entropic operator BB, defined by the curvature of its log multiplicity. A local Laplace expansion yields the fluctuation free energy Φfluc(M;B)=σξ22log⁡det⁡(M−1+B)+constΦ_{\mathrm{fluc}}(M;B)=\frac{σ_ξ^2}{2}\log\det(M^{-1}+B)+\mathrm{const}, analogous to an Occam factor. Under mild statistical conditions, this free energy is rotationally stationary exactly when [M,B]=0[M,B]=0, is minimized by pairing large eigenvalues of MM with small eigenvalues of BB, and generates a local restoring force against mismatch. Learning is therefore biased toward faster relaxation along entropically cheaper directions. This preference strengthens with training noise and vanishes in the deterministic limit, beyond gradient-flow accounts of operator alignment. For ReLU networks, we relate entropic curvature to the minimal rearrangement of activation boundaries required for a functional change and bound this structural cost by directional smoothness. Consequently, smooth directions are preferentially learned faster, in a data-adaptive manner, even as the learning operator evolves.
Sep 8, 2026cs.LG

When Does Scale-Invariant Optimization Become Unstable? An Exact Schedule Law with Weight Decay

Normalization renders large parts of neural networks effectively scale invariant, inducing a hidden feedback loop in which learning-rate schedules and weight decay interact through the parameter norm to control the effective step taken by the optimizer. We show that this interaction is governed by an exact discrete-time law: a single scalar quantity captures all schedule and decay forcing, while norm growth induces an opposing geometric self-quenching effect. This yields a sharp boundary that cleanly separates contraction- and expansion-dominated effective learning rate regimes. To understand the underlying mechanism, we provide exact analysis of a fully solved normalized regression model where the dynamics reduce to two dimensions and show that the balance point is intrinsically unstable, implying that constant learning rate with weight decay cannot stably maintain an interior equilibrium and instead produces recurrent behavior driven by discrete-time Jacobian structure. We further extend this perspective across optimizers through unified homogeneous-optimizer framework that reveals a structural dichotomy in self-quenching strength, providing a first-principles explanation for why adaptive methods exhibit systematically weaker stabilization under normalization. Across dynamical systems and neural networks (MLP, CNN, GPT2 / MNIST, CIFAR, wikiText, OpenWebText), the predicted law holds with high precision and enables direct control of training via the identified scalar, with performance peaking sharply at the predicted boundary. Together, these results isolate a single governing quantity for scale-invariant optimization, providing a precise and actionable lens on training dynamics, optimizer behavior, and schedule design in modern deep learning. Code is available in https://github.com/shasanamin/normalized-optimization-dynamics.
Sep 8, 2026cond-mat.stat-mech

Speed Limit for Information Acquisition in Stochastic Learning Dynamics

Neural networks acquire internal representations through learning. In this work, we formulate stochastic gradient descent (SGD) as a Markovian stochastic process and derive a Fisher-information flow speed limit that bounds the rate at which trainable parameters can acquire information about latent variables in the data-generating process. The resulting inequality decomposes the information flow into drift and noise contributions, thereby quantifying the roles of deterministic learning forces and SGD-induced fluctuations from an information-theoretic perspective. We verify the bound in analytically tractable basis-function linear regression, where the information budget predicted by the bound reproduces the ordering and characteristic time scales with which different latent variables are encoded in the learned parameters. These results establish Fisher-information speed limits as a quantitative framework for diagnosing when and how different aspects of the data-generating mechanism are acquired during stochastic learning.
Sep 7, 2026cs.LG

A Theoretical Analysis of Generalization Dynamics in Neural Networks under Gradient Descent with Weight Decay

Understanding generalization remains a central challenge in machine learning because it requires jointly considering data, architecture, and training dynamics. In this paper, we develop a theoretical framework that characterizes how these factors jointly shape generalization performance throughout training. More precisely, we study a broad class of neural networks trained under the ℓ2\ell^2 loss by gradient descent (GD) with weight decay, and prove the convergence of GD to a neighbourhood of the global minimizers of the empirical loss. By partitioning the space based on the input data, we then decompose the population error into data error, optimization error, and prediction variation error, and bound them separately. In particular, for the prediction variation error, which measures the oscillations of the learned function, we propose (local) approximate homogeneity and derive explicit cellwise and layerwise bounds for its evolution along the training trajectory. These bounds yield two important implications: a necessary condition of improved generalization explains differences in layerwise generalization behavior; a sufficient condition describes delayed generalization and provides a theoretical characterization of grokking.
Sep 3, 2026cs.LG

High-Dimensional Learning Dynamics of Attention-Indexed Models

Attention mechanisms are central to modern foundation models, yet their training dynamics remain poorly understood, especially when the attention matrices have extensive rank. In this work, we study attention-indexed models, a broad framework that can represent multi-layer and multi-head attention architectures. First, we show that, in a suitable high-dimensional limit, the population-loss landscape is characterized by a finite set of trace order parameters. In contrast, online stochastic gradient descent (SGD) is governed by an infinite hierarchy of matrix moments, which we show can be exponentially well-approximated by a finite truncated system. Second, this framework reveals that attention parameterization itself can act as an architectural implicit bias. Direct optimization of an attention matrix S∈Rd×dS\in\mathbb{R}^{d\times d} can remain trapped in an uninformative state. Tied attention (S=WW⊤S=WW^\top) induces an automatic symmetry-breaking mechanism and yields weak recovery in Θ(d2log⁡d)Θ(d^2\log d) samples. For untied attention, S=UV⊤S=UV^\top, we uncover a fast-slow mechanism: the pre-activation mean first evolves on a fast timescale, while the overlaps evolve on a slower one. Weak recovery on the Θ(d2log⁡d)Θ(d^2\log d) scale occurs when the state selected by the fast dynamics breaks the initial symmetry.
Sep 1, 2026cs.CL

Measuring Optimal Transport in Transformer Depth

A transformer carries each token's state from layer to layer, and the whole vocabulary carried together forms a cloud that moves with depth. We ask whether a trained network moves this cloud the way optimal transport would: at the cheapest cost, and along the map that pairs each token with its optimal destination. We measure both on Pythia-160m and Pythia-410m, with an exact assignment between consecutive layer clouds, a measured sampling floor, calibration on couplings known to be optimal, and a split of the cost into the common shift of the cloud and the token-specific moves. At the last layer, both models move their tokens where the optimal-transport map sends them, at the optimal cost for Pythia-410m and slightly above it for Pythia-160m. At the first layer they do not. In between, single layers can be judged on cost at only two of ten transitions, and blocks of several layers move the cloud at close to the optimal cost. The agreement at the last layer is much weaker at initialisation (0.64 against 0.86) and grows with training.
Aug 31, 2026cs.LG

How Temporal Correlations Shape Memory in Linear Recurrent Neural Networks

The linear recurrent neural network (LRNN) is a simple model for studying how much memory a network builds up as it trains. For uncorrelated inputs, earlier work found that training itself settles the network between keeping the past and reacting only to the present. Real sequences are correlated, and we solve the learning dynamics exactly for correlated inputs. In the solution, keeping the past carries a cost. The whole effect of correlation lands on that cost. This cost reduces to the earlier one when inputs are uncorrelated and grows once they are positively correlated. Three findings follow. (1) Correlation reshapes the course of learning, not only its end. Memory builds, overshoots, and is partly removed, and the settled network keeps less of the past. (2) Memory switches off at a threshold set by one number, how much each input resembles the one just before it. Neither sequence length nor longer-range correlation moves this threshold. Memory is worth keeping only when the task needs the previous input more than the current input already supplies it through correlation with the past. (3) The best network changes too. Zero error demands a feedthrough, a path that passes the current input straight to the network's output and remembers nothing, and training builds it unprompted when given one spare hidden dimension. Our work turns one property of the input into a prediction of whether a network learns memory and explains why correlated data turns recurrent networks into change detectors.
Aug 31, 2026cs.LG

Sparse Competition during Training For the Emergence of Specialized Modules

Modularity in deep neural networks has been proposed as a means of improving both interpretability and training by promoting disentangled representations and reducing redundancy. In this work, we study the emergence of modular structure through competition dynamics between groups of neurons during training. We introduce a method that (i) maintains near-baseline accuracy, (ii) induces usage-based modularity by sparsely routing inputs to neuron groups, and (iii) encourages specialization of these modules, such that their activations are correlated with input classes. We evaluate the proposed approach on ImageNet-100 and CIFAR-100 and show that with it, specialized modules emerge without module-level supervision. These modules capture a meaningful high-level structure in the data, with individual modules responding to semantic categories (e.g., dogs or vehicles). We also study the emergence of a hierarchical partition of sub-tasks depending on the number of modules. Our results suggest that competitive dynamics can serve as a simple mechanism for inducing functional modularity in standard architectures.
Aug 31, 2026cs.LG

Context Staircase: Signature-Aligned Dynamics of Token Embeddings under Small Initialization

Token embeddings are the basic representational units that connect discrete tokens with continuous computation in language models. Although modern language models learn embeddings from random initialization through gradient-based training, the dynamical mechanism by which meaningful embedding structures emerge remains unclear. In this work, we identify that the evolving embedding structures are closely related to token-conditioned label and contextual distributions, which we formalize as probability signatures. We observe a progressive learning process, which we term Context Staircase: embeddings learn the low-order statistic signatures of the data before the high-order ones. More specifically, we observe that early in training they align with the simplest, context-free signature linking a token to its label, and as training proceeds, they progressively reflect signatures involving more and more context tokens. We then analyze the gradient flow of embeddings under small initialization to explain this phenomenon, deriving embedding evolution equations for feed-forward and self-attention architectures. We further extend these observations to real language-model training. Finally, we show that these embedding structures play an important role in both task learning and the incorporation of semantic structure into the embedding space. Overall, our results provide a dynamic explanation of how data statistics and architecture jointly shape token embeddings in language models, and reveal an implicit bias in the space of data statistics: training proceeds from simpler, low-order statistical relations toward increasingly complex, context-dependent ones.
Aug 26, 2026cs.LG

Mapping the Emergence of Regularization-Driven Dynamics in Grokking

For overparameterized neural networks, many solutions can fit the training data equally well while behaving very differently on unseen samples. Grokking separates training fit from visible generalization, providing a window for studying how this selection develops during training. We sweep short, fixed-duration weight decay (WD) perturbations across the pre-generalization plateau and measure how they shift later generalization time. Across three grokking tasks, these shifts are unordered early in the plateau but later form a stable dose ordering before visible generalization, with stronger WD increases leading to earlier generalization and stronger WD decreases leading to later generalization. Test-loss barriers between perturbed and baseline generalization checkpoints collapse toward zero while the ordered timing effects persist. A similar response reorganization is observed under ℓ1\ell_1 regularization in the grokking setting of Junior et al. (2025). Drawing on Waddington's developmental landscape as an analogy, we call this combination of increasingly constrained solution selection and persistent dose-ordered timing shifts the canalization of grokking solution selection. Together, our response maps and loss-barrier measurements reveal a dynamical reorganization before visible generalization that is consistent with the theoretical picture of regularization-driven motion along a stable slow manifold (Boursier et al., 2025).
Aug 23, 2026cs.LG

Functional compatibility as a determinant of persistent neural learning

Neural networks can acquire new capabilities while damaging existing ones, but what determines whether new learning persists remains unclear. We identify functional compatibility, the extent to which incoming learning can coexist with behaviour that must be preserved, as an experimentally manipulable causal determinant of persistence. From identical neural states, we vary compatibility while matching unrestricted learning opportunity and imposing a common retention requirement. Persistent learning increases with compatibility across independent directions, convolutional and transformer architectures, vision and text, and a ten-seed replication. Learning rules and retention constraints determine how much compatible opportunity is retained, whereas nonlinear geometry limits the matched intervention at larger update norms. Functional compatibility therefore reframes stability-plasticity from preventing forgetting to determining which new learning can coexist with existing function and persist.
Aug 21, 2026cs.LG

Adam at the Edge of Stability: Adaptive Feedback, Provable Oscillation, and Gradient Reversal

The edge-of-stability (EoS) phenomenon of full-batch Adam has been widely observed, yet its underlying dynamical mechanism remains poorly understood. In this paper, we identify Adam's second-moment adaptation as a negative-feedback mechanism that drives the dynamics toward the stability boundary. We characterize this mechanism through the active curvature, namely, the preconditioned curvature along the preconditioned gradient direction, and establish rigorous characterizations in progressively richer settings: rank-one quadratics with momentum, diagonal quadratics, on which the active curvature separates from the sharpness, and general objectives. Importantly, the mechanism predicts gradient reversal of full-batch Adam near the edge: consecutive gradients repeatedly point in nearly opposite directions, as we observe across fully connected networks, ResNets, ViTs, LSTMs, GPT-2 medium, and Adam-family optimizers. Consistent with this picture, averaging iterates suppresses these fast oscillations and produces smoother and lower loss curves. Together, these results provide an important first step towards fully understanding the dynamical behavior of Adam's EoS through active curvature and gradient reversal.