Neural Network Training Dynamics
Momentum
41 papers in the last four weeks, up 273% on the four weeks before. 0.4% of all new papers.
Latest papers 349
We study whether persistent out-of-distribution (OOD) degradation can be predicted before it is directly observed using only source-side training dynamics. In a controlled shortcut-learning setting, a simple logistic regression predictor develops a clear prospective signal, while training time alone does not. Temporal summaries of the source-side quantities are substantially more informative than their current values. When transferred without additional training from a CNN to an MLP, confidence and entropy dynamics retain substantial predictive information. These results provide a proof of principle that source-side training dynamics can contain an early warning signal for future OOD failure.
Probability-Signature Dynamics: Unpacking Modular Addition Learning Within Two-Layer Networks
Neural networks trained on modular addition tasks often develop Fourier-structured representations that support exact generalization. While prior work has identified these Fourier circuits, the mechanism by which gradient-based training selects them from the data distribution remains unclear. We address this question using probability signatures, which express leading gradient interactions through conditional statistics of the training distribution. For modular addition, these signatures are cyclic shift operators and are diagonalized by the discrete Fourier transform, yielding approximately decoupled Fourier-mode dynamics. This explains the emergence of Fourier sparsity, frequency matching, and phase alignment. The same framework resolves a puzzle under label noise: corrupted examples can show faster early loss decrease than clean examples, despite lacking a coherent generalization rule. We show that noise increases conditional label collisions, strengthening early shared-coordinate reinforcement. Finally, this method can be applied to other operators. Taking XOR as an example, we observed the predicted frequency in experiments.
Rare Gate Disagreements Can Limit Plasticity: When Gradient Flow Mispredicts Finite-Batch SGD
Population gradient flow is a common tool for reasoning about how neural networks adapt, including after pretraining. We show that it can mispredict finite-batch stochastic gradient descent (SGD) qualitatively, and we trace the discrepancy to a specific mechanism. In a two-unit ReLU regression, a source task drives the two neurons toward positive proportionality and a target task rewards separating them. After source training for time , gradient flow recovers on the target in time linear in . Online SGD with batch size and step size in both phases instead fails with high probability throughout a horizon of order once , uniformly on an explicit set of initializations with Gaussian probability above one percent. For each fixed , small-step SGD still recovers, so the failure requires the joint limit of small steps and long pretraining. At the target clone, the population instability is carried entirely by inputs on which the two ReLU gates disagree. For units at angle these inputs form a wedge of probability , and weight decay shrinks the angle exponentially during pretraining. On every other input both units receive the same random linear update, which contracts their separation in conditional expectation. Bounding the cumulative probability of sampling the wedge along the exact online recursion, without a diffusion approximation, shows that recovery with fixed probability from an identical source-gradient-flow checkpoint, within updates, requires target samples and batch size , where counts updates and is the weight decay. In simulations, recovery is approximately a function of the disagreement budget and saturates in the horizon.
Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals
On-policy distillation (OPD) enables effective capability transfer between language models, yet the mechanisms underlying its failures are not fully understood. Across code generation and mathematical reasoning, OPD with larger-scale teachers exhibits early loss plateaus, with an average final loss reduction of 25.1% after 200 updates, compared with 96.2% for self-RL teachers, obtained by further reinforcement learning (RL) training of the initial student. To understand this difference, we analyze OPD as an idealized continuous-time dynamical system in the small-learning-rate limit. Our training-log diagnostics associate these plateaus with an early decline in a gradient-based learning-signal proxy while substantial loss remains; these measurements do not establish why the underlying gradient weakens. We further prove a local recovery guarantee for teachers sufficiently close to the initial student in a shared parameterization under regularity conditions, offering a conditional explanation for the success of self-RL teachers in our experiments. Across runs with and without loss plateaus, we observe small relative parameter changes (0.025-0.098%) and high similarity between the student's representations before and after OPD (linear CKA across layers). These observations suggest that limited representation adaptation may contribute to learning-signal collapse, a hypothesis that remains to be tested. Code is available at https://github.com/leizhao7/opd-learning-signals.
Fluctuations of Nonlinear Observables in Mean Field Neural Network Training
Mean field limits describe the training dynamics of wide neural networks through the evolution of the empirical distribution of their parameters. Although functional central limit theorems characterize the asymptotic fluctuations of this distribution, quantities of practical interest are typically nonlinear observables of the parameter distribution rather than the distribution itself. In this work, we show how these mean field fluctuations propagate to finite dimensional nonlinear observables for shallow neural networks trained by stochastic gradient descent. Working in the weighted Sobolev space in which the limiting fluctuation process is constructed, we apply a functional Delta method under ordinary Fr{é}chet differentiability, without requiring Lions derivatives with respect to the measure variable. We obtain a central limit theorem for the observables and, under a suitable representation of their differentials, an explicit covariance formula inherited from the underlying mean field fluctuation theory. We also study whether prescribed quantities of interest can be recovered from the selected observations. Under a constant rank assumption, we prove that a quantity of interest factors locally through the observation functional if and only if, throughout a neighborhood, the kernel of the differential of the observation is contained in that of the quantity of interest. Thus, a differential condition expressed directly in the ambient Sobolev space yields an exact nonlinear local factorization. These results provide a framework both for quantifying finite-width uncertainty on observable, statistically or physically meaningful quantities and for assessing whether the chosen observations contain the information required to identify them.
When does a network's training history predict its future learning better than its current state? Evidence from a response probe and a forecasting screen
Networks that behave alike now can still learn differently when training continues. Work on loss of plasticity and critical periods shows that the path to a state shapes what follows; it does not show whether the path carries information that a measurement of the state itself misses. We ask when the training history of a network predicts its future learning better than its current state. In a main study, small multilayer perceptrons were trained under three history regimes (42 histories), and future learning was measured at four checkpoints by a short probe: a copy of the network trained for 100 updates on a new task. Before the prediction result was read, the protocol checked the probe. It responded monotonically to a function-preserving rescaling of hidden units, repeated measurements agreed (intraclass correlation 0.940, [0.903, 0.997], in the least reliable class, mean of three repeats), and a re-initialisation of units was visible directly after it but not 100 to 200 updates later. A history state of at most four dimensions did not improve on a calibrated model of the current state (gain -21.4%, 90% interval [-91.9, 8.1]; required in advance: 10%). A companion screen on 1,560 synthetic regression runs asked the same question for a target further away, the final error of the run. There, history models forecast better than the current validation error after 12 of up to 240 epochs (compact state 30.3%, [15.8, 39.4], a contextual comparison) and were not distinguishable from it after 48. In both studies the history was informative only while the current state was not yet informative about the target; this reading was formed after the results.
Stability and Diversity of Networked Self-Consuming Generative Ecosystems
The widespread deployment of generative AI has made it increasingly difficult to distinguish synthetic content from real data. Consequently, synthetic data is inevitably incorporated into the training pipelines of future model generations, forming a self-consuming training loop. Prior work has studied the effects of such recursive self-consuming training, but analyses have largely been limited to isolated models, where a model consumes only its own synthetic data, or to simplified interactions between two models. This paper takes a first step toward understanding networked self-consuming generative models, in which multiple models consume synthetic data generated by one another through complex interaction pathways. We introduce a theoretical framework representing models as nodes in a directed, weighted graph, with edge weights governing the flow of synthetic data among models. Using this framework, we analyze the long-term behavior of networked models under retraining dynamics, establishing conditions for convergence and characterizing the resulting fixed points. We further investigate how the system's long-term stability and diversity are shaped by each model's access to real data, cross-model data consumption, and the structure of the interaction graph.
Global Exponential Convergence of Two-Layer Linear Network Training
We prove global exponential (linear) convergence with an explicit rate in the rich scaling for wide two-layer linear networks trained with smooth Polyak-Lojasiewicz predictor losses. Gradient flow in the factors closes exactly in terms of a finite-dimensional Bures flow of the neuron law covariance, in which the predictor dynamics are preconditioned by hidden covariance blocks. Mean-field conservation laws provide uniform spectral lower bounds on the hidden preconditioning blocks when the initial covariance satisfies a spectral support gap condition. This condition encompasses positive definiteness while still allowing for singular initializations. For an initial covariance , the loss converges to the global minimum with linear rate at least , where is the PL constant. We establish stability of this rate under finite-width sampling, as well as global convergence of factor gradient descent for an explicit stepsize interval depending on smoothness, the initial loss, and conserved spectral margins. Our argument extends layerwise to deep linear ResNets, subject to a residual-path bound. In the case of heavy-ball momentum, training dynamics close instead over positions and velocities in terms of a lifted phase covariance. Linear convergence holds under an explicit condition on the energy and damping, specifying a window of admissible dampings. For two-scale white initializations, this interval is nonempty for sufficiently large position scales, with a fixed initial loss gap and velocity covariance. Numerical experiments illustrate the covariance geometry and compare the predicted and observed rates.
Quadratic Weak-to-Strong Generalization in Random Feature Networks via Random Matrix Theory
Weak-to-strong generalization is the phenomenon where a strong student model trained with labels produced by a weak teacher model is able to generalize better than the teacher. In this paper, we study this phenomenon in two-layer random feature networks where the model strength is determined by its width. Using tools from random matrix theory, we derive deterministic equivalents for the population errors of an optimally trained teacher and a student trained with gradient flow. For ReLU activation and a pure spherical harmonic target, we obtain sharp asymptotics under a Gaussian universality assumption, showing a quadratic improvement: the student error scales as the square of the teacher error. These results attain the general lower bound of Medvedev at al (2025). We also analyze how the student behaves under more general stopping times and targets supported on multiple harmonic degrees, characterizing the regimes in which weak-to-strong generalization occurs and identifying the transition between quadratic, non-quadratic, and no improvement.
Neural Fields Encode Adaptation Geometry
Neural fields are usually evaluated by how well they reconstruct an observation. We show that this misses two useful properties of a fitted network: how easily it can adapt to new observations, and what its weights retain from earlier ones. We study these properties as adaptation geometry. For images, we meta-learn class-specific initializations, adapt each one to a new image, and measure how much the network must change to fit it. A simple local linear model closely predicts this adaptation cost, while replacing one network's tangent kernel with another's substantially worsens the prediction. Adaptation thus depends on the local geometry of the fitted network, not only on its current reconstruction. For physical fields, we repeatedly fit the same network to observations from a sequence. Its weights then retain information about that history. When two wave histories end at exactly the same observation, the final weights recover the sign of the wave velocity with 68.6% accuracy, whereas the current observation alone contains no such information and gives 50%. These two phenomena are quantitatively linked: tangent-kernel eigenvalues predict both which changes are easy to learn and how quickly they are overwritten by later fitting. Together, these results show that neural fields contain useful information beyond what they currently reconstruct: in how they can change and in how they got there.
Mind the Drift: Diagonal Linear Networks Under Large Learning Rates
Large learning rates can qualitatively change the trajectory of neural network training, often pushing optimization into regimes far from classical gradient-flow behavior. The Edge of Stability (EoS) offers a valuable lens on the dynamics such learning rates induce. We study corresponding dynamics in diagonal linear networks, where we uncover a competition between two distinct implicit biases that jointly determine the sparsity of the recovered solution in regression settings. Complementary to the Gain, which captures the average discretization error accumulated by Gradient Descent relative to Gradient Flow, we derive a closely associated but overlooked quantity: the Drift. Under large learning rates, it describes an imbalance between different discretization errors and represents a systematic shift in the optimization trajectory. While the Gain grows monotonically in certain regimes, and can bias towards denser, flatter interpolators, the impact of the Drift depends on its alignment with potential solutions, which can either counteract or reinforce the effect of the Gain. Consequently, its behavior drives model selection, particularly during early training epochs. To validate our theoretical insights, we introduce an intervention that actively steers the Gain to recover sharper, sparser solutions. Thus, our analysis reveals that large learning rates do not universally hinder the recovery of sparse solutions. On the contrary, they can be harnessed to control the implicit bias of training.
Generalization in Neural Networks Through the Lens of Magnitude Potential
Explaining generalization and training dynamics in neural networks remains a challenge, and various approaches have been developed to study different aspects of these phenomena. In this paper, we introduce the idea of {\em magnitude potential} -- a quantity based on the theory of metric magnitude -- that reflects how well an arbitrary point is represented by a given set. We find that this basic quantity can be applied to examine various features in neural generalization. The ratio between the magnitude potential with respect to a class and with respect to the entire data, computed at the logit layer, is informative of the representation of the point. In experiments, these ratios for individual training points are found to be correlated with the Feldman memorization scores. Magnitude potential ratios aggregated across points detect structural changes in the decision boundaries and provide a geometric indicator of grokking in modular arithmetic. Although the magnitude potential ratio and neural collapse are both closely associated with intra-class and inter-class geometric structure, the magnitude potential ratio remains informative even when neural collapse is explicitly suppressed.
Why Does Train-Validation Separation Emerge? Update-Pressure Density Dynamics in Pretrained Backbones
Train-validation separation is the evolving difference between performance on observed training examples and a finite held-out validation set. We propose a dynamic structural account of how this gap develops during adaptation of pretrained models: continued fitting can shift update demand from broadly reusable support toward narrower support with weaker held-out transfer. A conditional local model links this shift to increasing heterogeneity in gradient allocation and train-validation separation. Fixed training probes make this structural evolution observable without validation examples entering the readouts; held-out performance is used separately to evaluate its relation to the gap. In a constructed hierarchy implemented with a residual multilayer perceptron (ResMLP), increasing the target share of example-private features from to to , while preserving the relative mixture among the four shared feature levels, increases the final mean accuracy gap from to to across five runs per condition. Masked-input losses measured separately on training and validation examples expose the corresponding transfer asymmetry. The natural language processing (NLP) analysis uses 10-epoch runs of RoBERTa, DeBERTa, and Qwen on six datasets (90 runs): the training-probe-weighted within-class and overall dispersion readouts each have positive raw and smoothed level correlations with the accuracy gap in all 90 runs. Raw changes paired at approximately one-epoch intervals remain positively associated in 86/90 and 87/90 runs, respectively. A 40-epoch ResNet-18 study tests both readouts on three vision datasets. Together, controlled simulation, NLP, and vision support the dynamic structural account across settings, with real-model evidence testing its observable predictions under the specified monitors.
Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning
End-to-end backpropagation has been the dominant mode of training in deep learning, allowing for the coordination of parameter updates across layers of a neural network. Prior studies have explored alternative -- and, in some cases, simpler -- training mechanisms, showing that they can sometimes achieve performance similar to backpropagation. However, the architectural conditions under which locally optimized networks, which avoid end-to-end backpropagation of error, can learn representations comparable to those learned through end-to-end training remain unclear. We aim to answer this question in the context of self-supervised learning, an important framework for large-scale pretraining in artificial intelligence. Here, we investigate how network width and depth affect the efficacy of greedy layer-wise and end-to-end self-supervised training in convolutional networks. We find that in wider networks, the benefits of end-to-end backpropagation over greedy layer-wise training shrink: in relatively shallow and very wide networks, we even observed higher performance in models trained with greedy layer-wise training. Subsequent analysis of the representations formed by these networks shows that very wide greedy-trained networks exhibit more favorable representational geometry than do networks trained end-to-end with backpropagation. This work shows that width can compensate for restricted credit assignment and identifies differences in representational geometry as a potential mechanism for their improved performance.
Disentangling Computation in Multi-Task Neural Networks with the Green's Operator
How is computation organized and reused across tasks and time in a trained recurrent network? Most analyses emphasize the geometry of neural activity, dynamical motifs, or local perturbation growth. We instead study the network's global first-order perturbation response. The finite-horizon Green's operator maps perturbations at each source along a trajectory to their downstream state-space responses and therefore directly represents perturbation routing. Simple reductions of this operator provide task-to-task and time-to-time views of the same computation, while matrix-free products make these views accessible without constructing the full operator. In a flexible multitask recurrent network, task reductions reveal structured reuse of known computational motifs, while temporal reductions reveal causal pathways and how they emerge during training. Our main point is simple: the Green's operator provides a global response geometry for mapping the organization of learned dynamical computation.
The Life Cycle of a Massive Activation: Stochastic Birth, Weight-Decay-Driven Growth, and Competitive Consolidation
Massive activations, residual-stream coordinates with magnitudes far larger than typical activations, are associated with attention sinks in transformers, but how their scale is regulated during training remains incompletely understood. Combining training-trajectory analyses and controlled interventions, we trace their emergence, growth, and consolidation. Sink-carrying channels vary across random seeds but stabilize early within each run. Over longer training, surrounding channels erode and the sink concentrates onto a few redundant carriers. Across ablations, gradient attenuation follows the sink token's collective root-mean-square magnitude rather than any single channel, making collective scale central to understanding their effects. Our central result is that weight decay causally controls the turnover of global activation scale. In controlled continuations, removing decay near the peak allows this scale to keep rising, whereas retaining it produces decline even at constant learning rate. We develop a balance model for the rise and peak of massive-activation magnitude, in which AdamW-preconditioned growth opposes weight decay. Sweeping the decay coefficient shifts peak timing approximately log-linearly and yields peak magnitudes scaling approximately as , consistent with this balance. Optimizer measurements further show that preconditioning sustains the large-channel cohort against decay even when raw maintaining forces are too small to do so. Together, these findings connect the observed life cycle to scale-regulating training dynamics and establish weight decay as a training-time lever on activation magnitude.
How Does Local Landscape Geometry Evolve in Language Model Pre-Training?
The scale and expense of pre-training language models make efficient hyperparameter tuning essential, yet a principled guidance is still missing. In this work, we analyze language model pre-training dynamics from a local landscape geometry perspective. Our study reveals two distinct phases. In Phase I, sharpness of the local landscape is initially high, leading to instability and loss plateaus under large learning rates (LRs). The landscape shifts from sharp to flatter regions early in training. This dynamic explains the necessity of LR warmup and further suggests that larger peak LRs require proportionally longer warmup periods. In Phase II, the local landscape is governed by the gradient noise scale. Our theory identifies a depth flatness trade-off: high noise from smaller batches widens the loss basin, whereas reduced noise from larger batches deepens it. This theory motivates a dynamic batch-size (BS) scheduler that begins with a small BS and increases it late in training. Together, we provide a unified view of loss landscape evolution, which translates into actionable tuning strategies for large-scale pre-training.
Stable Transformers for Graph Generation
Graph generative models increasingly rely on Graph Transformers (GT) to capture complex dependencies among nodes and edges. While deeper architectures should provide greater expressive capacity and a broader receptive field, their effectiveness can decline with depth: repeated self-attention progressively contracts node representations, impeding information flow and gradient propagation. We analyse this phenomenon from a dynamical systems perspective, focusing on how the denoiser's spectral dynamics affect graph generation. We show that standard GT denoisers become increasingly dissipative as depth grows, leading to vanishing gradients and representation collapse. To isolate the effect of these dynamics, we construct a permutation-equivariant GT with inherently stable, non-dissipative transport. We also introduce a damping mechanism that continuously interpolates between non-dissipative and increasingly contractive regimes, enabling a direct assessment of how dissipation influences generation. Experiments on synthetic and molecular graph generation benchmarks show that the gap between these regimes widens with depth: non-dissipative dynamics preserve representation diversity and gradient flow, sustaining strong generative performance, whereas greater contraction progressively impairs it. These findings identify the denoiser's dynamical regime as a key design factor for deep graph generative models.
The Conflict Between Logic and Memory: Learning Higher-Order Interactions in Shallow MLPs
A network can fit its training examples while failing to recover the rule that generated their labels. We examine this separation in single-hidden-layer multilayer perceptrons (MLPs), using synthetic tasks that control interaction order and the presence of nuisance inputs. We establish elementary benchmark properties: pure parity contains no predictive lower-order marginals, admits an exact Bayes posterior, and can be represented on clean latent inputs by a width- ReLU network. Experiments then identify distinct optimization outcomes. In a matched order-2--4 sweep, SGD, Adam, and Muon all reach 100% peak test accuracy at order two; at order three they reach 96.25%, 50.87%, and 76.82%, respectively, while Muon reaches 99.21% at order four. In a separate mixed-order task, freezing only the first-layer weights connected to independent nuisance inputs raises AdamW's epoch-10 accuracy from 44.73% to 95.07%. Removing the same inputs only at test time raises it to 48.38%. Thus, nuisance-weight learning changes the training outcome beyond its immediate effect on prediction. Bias interventions expose a connection between target symmetry and shallow ReLU representations. In a compact signal-only regime, both SGD and Muon learn orders five through eight, with higher SGD peak accuracy at orders nine through eleven. Together, the results show how optimization and nuisance learning constrain the higher-order rules realized by a shallow network.
Awakening of the Buddha: Subspace Learning During Population-Loss Plateaus
Population loss can remain nearly constant while a neural network learns a substantially more predictive representation. We establish this separation for two-layer ReLU and leaky-ReLU networks trained on Gaussian inputs by simultaneous fixed-step population gradient descent on all parameters. For structured additive teachers whose links are positive mixtures of Gaussian-damped cubics in , we give explicit conditions under which small IID Gaussian initialization yields a high-probability guarantee: at a checkpoint during a high-loss plateau, minimum alignment between the rank- teacher subspace and the leading -dimensional eigenspace of the predictor's average gradient outer product (AGOP) increases by at least , and the minimum refit MSE under unchanged coefficient budgets decreases by more than , both relative to initialization. The same trajectory subsequently attains a trained loss below every value in the plateau window. A complementary result treats unequal-weight cubic teachers and small additive Sobolev perturbations using projected-feature refits. For SwiGLU networks with an exactly fitted intercept, we prove leading-AGOP alignment during a loss plateau at fixed width and dimension as Gaussian initialization vanishes, for square-integrable teachers with nonzero Hermite content of degree one, two, or three. A rank-one cubic specialization also gives simultaneous unrestricted-refit gains at a prescribed width. An approximation lower bound further shows that certain interaction targets retain nonzero error when ridge neurons are restricted to shared orthogonal axes within the teacher subspace. Population-moment experiments with ReLU students across 21 teachers and 50 initializations per teacher complement the analysis.
A Dynamical Theory of LoRA in Continual Learning
Despite the widespread use of Low-Rank Adaptation (LoRA), little is known about its dynamics in continual learning and the mechanisms by which low-rank updates affect catastrophic forgetting. We provide an asymptotically exact dynamical characterization of LoRA in a solvable two-task teacher-student model. In the high-dimensional online-learning limit, we derive a closed system of ordinary differential equations for a finite set of macroscopic order parameters, yielding exact expressions for the generalization errors throughout both the initial Task 1 learning phase and the subsequent LoRA fine-tuning on Task 2. The theory quantitatively matches finite-dimensional simulations and exposes two characteristic effects of LoRA: low-rank adaptation reduces interference with features learned on the first task, but its initialization slows adaptation to the second task. Building on this mechanistic picture, we analyze a state-dependent masking strategy that freezes hidden units carrying the strongest first-task representations and restricts adaptation to the complementary subspace. This structural partitioning markedly reduces forgetting, while preserving plasticity on the new task. Our framework further clarifies the role of adapter rank: transfer improves only up to the intrinsic dimensionality of the target task and saturates beyond it, while forgetting continues to grow with rank. These results provide a dynamical and geometric account of how low-rank adaptation organizes information across sequential tasks and are qualitatively reproduced on a sequential MNIST benchmark.
Learning Under Forgetting: Statistical Support-Selective Retention in Stochastic Training Dynamics
Prior work has shown that neural networks exhibit implicit biases toward low-complexity structure (e.g., spectral bias), memorization dynamics, and compression-like effects during training, but a unified dynamical account of selective retention remains incomplete. We propose Repeated Reinforcement with Persistent Forgetting (RPF) dynamics, a minimal framework in which repeated exposure reinforces patterns and structures that recur in the data, while persistent forgetting attenuates learned information. This view treats forgetting not merely as a failure mode, but as a selection mechanism. We build the theory in three successive layers. First, in an independent-feature model, we derive an exposure-selective survival law and a support-dependent retention boundary characterizing which patterns persist under forgetting. Second, in a shared-parameter model, we show that forgetting induces spectral filtering over covariance modes, preserving strongly supported shared components while suppressing weak ones. Third, under small-step and norm/coding approximations, we show how RPF dynamics induce an implicit trade-off between data fitting and the cost of stored information, yielding Minimum Description Length (MDL)-like compression. Controlled experiments provide evidence for this reinforcement--forgetting selection mechanism in scalar memories and a nonlinear shared network. Joint reinforcement and attenuation interventions shift conditional retention, while matched exposure counts reveal forgetting-dependent effects of reinforcement timing and changes in the composition of the retained set. Together, these results show that repeated reinforcement and persistent forgetting jointly provide a controllable source of inductive bias beyond neural architecture and scale.
Where Does Randomness Matter in Neural Cellular Automata?
Stochastic cell updates are often used throughout the life of a neural cellular automaton (NCA), from backpropagation through time to final rollout. This leaves two questions entangled: does update randomness help learn a useful rule, and must that randomness remain at execution? We separate training and evaluation update modes in controlled Growing NCA experiments, then vary the states shown during training. Under the standard constant-rate persist recipe, asynchronous training passes the short-horizon quality test in 10/10 runs, compared with 3/10 synchronous runs. All ten asynchronous models also retain the target for 4,096 steps under deterministic evaluation. For a scalar translation-invariant lattice, we derive an exact mean-square criterion: random masking can damp mean modes, but it also injects variance, and a mean-only test misclassifies four non-marginal settings. Finally, among 30 models that all pass the same reconstruction test, eight of ten grow-trained models become off-target at 4,096 steps, while all persist and regenerate models retain the target; damage recovery separates persist from regenerate. The results distinguish optimization reliability, execution mode, and task-specific behavior instead of treating them as one stability property.
The Hidden Ratio in Adam: Stable Structure, Compression, and Sign Dynamics
Adam is the default optimizer for training modern deep neural networks, yet its adaptive behavior remains poorly understood due to the complex interaction between its first- and second-moment exponential moving averages (EMAs). We study Adam in the tied- regime, where the two EMA decay rates are equal, and show that its adaptive dynamics can be expressed through a transformed ratio with approximately scale-stable behavior. Empirically, this transformed ratio exhibits a stable, heavy-tailed distribution across tasks, model scales, and training stages, in contrast to the variability of raw moment magnitudes. This empirical stability has both practical and conceptual consequences. First, we derive a recurrence for the transformed ratio, yielding a reparameterization of Adam that replaces the second moment with a compressible state. Leveraging its stable distribution, we show that a fixed 4-bit codebook is sufficient in our experiments to store this state without auxiliary scaling, achieving performance competitive with full-precision Adam. Second, the transformed ratio view clarifies Adam's connection to sign-based methods: Adam reduces to sign-based momentum modulated by the transformed ratio, and replacing it with a constant recovers Signum as a limiting case. This perspective further provides a simple rule for transferring learning rates between the two methods. Together, these results suggest that tied- Adam admits a simple and approximately stable ratio structure underlying its adaptive behavior and demonstrate its utility for both analysis and efficient implementation.
First Learn, Then Memorize: The Spectral Bias of Diffusion Models
Diffusion models trained on a finite dataset first learn to generate novel, high-quality samples and only much later collapse onto their training set. We identify the mechanism behind this separation of timescales and the object that probes it. The training dynamics of the score function are governed---exactly, and at any width---by the Gram matrix of the Neural Tangent Kernel (NTK) evaluated on the noisy training data, so the timescales of generalization and of memorization must be encoded in its spectrum. We show that they are, and that the structure responsible has no analogue in standard kernel settings. The use of multiple noise realizations per sample ( noised copies at a fixed noise level) in the score-matching loss is what restructures the Gram matrix spectrum into two distinct parts. The first, of large eigenvalues, carries the global features of the target distribution and is present already for . The second, which the repeated noising creates, consists of the smallest eigenvalues and is supported on eigenvectors aligned with the sample-specific noise directions; it sets a memorization timescale parametrically larger in the training set size . We establish this picture on two fronts. Analytically, we solve the spectrum in the lazy high-dimensional limit for both linear () and polynomial () sample complexities, and prove through a bias--variance decomposition that the first bulk minimizes the approximation error while the second drives the error associated with memorization. Empirically, we show the same two-bulk structure in Convolutional NTKs on CelebA and in finite-width U-Nets trained well beyond the lazy regime, and we make the link causal: truncating the Gram matrix at rank tunes the generalization--memorization transition, and an penalty targeting the second bulk suppresses memorization in feature-learning U-Nets.
Weighting Schedules Govern What and When Score-Based Generative Models Learn from Multimodal Data
Score-based generative models generate new samples by integrating a time-dependent drift that carries Gaussian noise onto the target distribution. In practice this drift is modeled by a neural network, trained on a loss integrated over time with a weighting schedule . Along the backward dynamics, and for multi-modal distributions, trajectories commit to modes of the target within a narrow time window, the \textit{speciation time}. In this work, focusing on high-dimensional data, we decompose the integrated loss into its single-time contributions and analyze each at fixed signal-to-noise ratio : we show that sets the rate at which each feature of a multimodal target - the mode directions and their relative weights - is acquired during training. Crucially, at high all mode directions are acquired together, on a single timescale insensitive to their amplitudes, while the relative weights are not learned at all. Only near the speciation time, where becomes of order one, do all features become learnable, each on its own timescale: the weights are acquired jointly with the directions, and the directions at rates set by their relative amplitudes. For models trained on time-integrated objectives, the learning dynamics is then governed by how much of the weighting effectively sits near the speciation time, which provides insights on design choices. These results follow from an exact high-dimensional analysis of the training dynamics of unbalanced and hierarchical Gaussian mixtures. Numerical experiments on image and human genome haplotype generation recover the predicted hierarchy of learning timescales in more complex settings.
Fast Learning Rate Transfer in Shallow Linear Networks at Growing Training Horizons
Hyperparameter transfer across model width can substantially reduce the cost of tuning large neural networks, but its behavior when the training horizon grows with width is not fully understood. Building on the framework of fast hyperparameter transfer (Ghosh et al., 2026), which formalizes when transfer is effective, we investigate conditions that ensure fast transfer in the growing-horizon regime. Specifically, we study learning-rate transfer in a shallow linear network with a single trainable hidden matrix, trained by full-batch gradient descent. Under additional spectral assumptions, our main results are threefold. (i) We prove fast learning-rate transfer as whenever . (ii) We characterize the transfer rates through the finite-width perturbation scale, the first-order sensitivities of the loss and its learning-rate derivative to finite-width perturbations, and the local loss curvature. (iii) We derive limiting distributions for the optimal learning rate and optimized loss, governed by fluctuations associated with the extreme eigenvalues of the data Gram matrix. These results clarify how spectral structure and local loss sensitivities govern learning-rate transfer at growing horizons.
Muon Sublates the Edge of Stability in LLM Pretraining
Muon is increasingly used for language-model pretraining, yet its large-step dynamics are not captured by the classical edge-of-stability (EoS) picture of gradient descent (GD). In GD, loss neutrality, equal-magnitude update reversal, and marginal stability meet at a single learning-rate-dependent edge. We show that Muon breaks this coupling. For stochastic no-momentum Muon, we derive a coherence-corrected conditional loss-neutral boundary , while temporal alignment follows a separate geometry. Controlled experiments show that loss balance and temporal alignment respond differently to learning rate and batch size. Across our language model experiments, the 130M Llama-like LLM runs exhibit loss-boundary tracking with weak negative alignment, whereas the studied 1B LLM configuration shows stronger partial cancellation; in both settings, directions remain far from coherent reversal while training continues to improve. These results support a split EoS picture for Muon: a stochastic loss-neutral edge survives, but it is not accompanied by a universal temporal-direction signature. The source code for reproducing the experiments can be found in https://github.com/cyzebra/Muon-Sublates-the-Edge-of-Stability-in-LLM-Pretraining
When Sparse Reward Meets Dense Distillation: Training Dynamics of On-Policy Distillation
Reinforcement learning with verifiable rewards provides a sparse post-training signal: a single binary outcome evaluates the entire rollout, and every token receives the same sequence-level advantage regardless of its individual contribution. To complement this sparse supervision, a growing family of methods adds a scalar-weighted teacher KL term to the policy-gradient objective, providing dense token-level guidance that may be unreliable at some positions. Despite the benefits of combining these signals, their interaction during optimization can destabilize joint training. To understand how this instability develops, we study the learning dynamics of hybrid reward--distillation training through a neural tangent kernel (NTK) analysis. We introduce the cross-signal NTK , a token-level statistic that measures the alignment between reward and distillation gradients at position n. Through this analysis, we identify two failure modes: 1 Magnitude drowning, where the reward gradient exceeds the distillation gradient by orders of magnitude, so that even weak directional conflict can cause the distillation loss to rise despite its explicit inclusion in the training objective; and 2 Localized directional conflict, where the sequence-level advantage and the teacher's position-specific distribution induce opposing updates at the same token (). The severity of these effects depends on the optimization regime: the gradient-norm ratio varies by roughly an order of magnitude across tasks, and our experiments reveal an empirical threshold beyond which naive mixing can lead to persistent training collapse. Motivated by these findings, we introduce the M3 family, which combines magnitude normalization with three strategies...
Two-Timescale Fine-tuning Provably Learns New Features for Two-Layer ReLU Networks
Fine-tuning pre-trained models on specialized tasks with scarce data is central to modern deep learning. Despite its empirical success, theoretical understanding of fine-tuning remains limited. We introduce a Gaussian multi-index setting to study fine-tuning from pre-trained weights, where the teacher network has features, of which are learned during pre-training and one of which must be learned during fine-tuning. For two-layer ReLU networks, we show that two-timescale training, i.e., updating the outer weights infinitely faster than the hidden ones, learns the new task-specific feature while preserving the pre-trained ones in the model representation. Moreover, only fine-tuning samples are required for this recovery, independently of the number of pre-trained features. In contrast, with random initialization, the same number of samples is insufficient to recover the target parameters. Our results therefore demonstrate that pre-training can induce an implicit bias with a clear statistical advantage over random initialization, enabling feature learning from scarce fine-tuning data.