cs.LGJun 9, 2026

Conservation Laws from Data Symmetry in Neural Networks

Authors: Jakob GalleyVahid ShahverdiAxel Flinth

Organizations: Umeå University, Umeå, Sweden

Abstract

We explore whether intrinsic symmetries of the training data lead to conserved quantities during gradient-flow training of neural networks. Under the assumption that the loss function is analytic and non-polynomial, we prove that data symmetries generically do not induce any additional integrals of motion. For mean squared error (MSE) loss, on the other hand, there are situations in which data augmentation yields extra conserved quantities. We build a framework, utilizing \emph{tensorizable networks} to describe this phenomenon. Tensorizable networks are a family of architectures whose dependence on parameters and inputs can be separated using an intermediate representation. They include linear and polynomial networks, as well as Lightning Attention.

Explore similar work

May 16, 2026cs.LG

Prediction Is Not Physics: Learning and Evaluating Conserved Quantities in Neural Simulators

A diffusion model trained on Hamiltonian trajectories can achieve rollout MSE near 10310^{-3}, but the standard deviation of its energy over time is between 7500 and 36000 times larger than the ground-truth energy standard deviation, indicating a failure to preserve conservation laws. This gap motivates our central question of whether neural networks can learn or select globally conserved quantities from physical trajectories. We investigate this across three Hamiltonian systems: projectile motion, pendulum, and spring-mass. We use a structured T(v)+V(q)T(v)+V(q) energy model, a black-box Conservation Discovery Network (CDN), a polynomial CDN, and a conditional diffusion baseline. The structured network reaches R20.9999R^2 \geq 0.9999 against analytical energy on clean data, while the black-box CDN reaches R20.996R^2 \geq 0.996 when trained with temporal consistency plus a small alignment loss to analytical energy at t=0t=0 (λalign=0.2λ_{\mathrm{align}}=0.2). With λalign=0λ_{\mathrm{align}}=0, CDN Pearson R2R^2 collapses on pendulum and spring-mass (<103< 10^{-3}), showing that temporal consistency alone is not enough to reliably identify the true energy. Under 1%1\% additive Gaussian noise, the CDN outperforms the structured model on the projectile and spring-mass systems, suggesting that the CDN may be more robust to noisy inputs in this setting. However, the polynomial CDN is sensitive to training configuration: it achieves R2=0.78R^2=0.78 under a short training schedule on the pendulum system, but reaches R2=0.9998R^2=0.9998 with more training time and data, regardless of whether noise is added.
Andrew Bukowski, Aditya Kothari, Simba Shi +1
Jun 16, 2026cs.LG

Conservation Laws for Modern Neural Architectures

Understanding gradient descent dynamics is key to explaining the success of over-parameterized models, where implicit bias manifests through conservation laws in gradient flow. While such laws are well understood for linear and ReLU networks, they remain largely unexplored for modern architectures. This work develops a unified framework to characterize conservation laws for contemporary models, including feedforward networks with GELU, SiLU, and SwiGLU activations, multihead attention with sinusoidal and rotary positional encodings, and Mixture-of-Experts architectures under diverse gating designs. Our theoretical findings are supported by experiments that validate the predicted invariants.
Viet-Hoang Tran, Vinh Khanh Bui, Tan Lai Ngoc +3
Jun 16, 2026cs.LG

A Link between Shock-wave Theory and Symmetry-reduced Stochastic Gradient Descent for Artificial Neural Networks

We develop a mathematically explicit link between shock-wave theory and the symmetry-quotiented learning dynamics of stochastic gradient descent, drawing on differential geometry, Lie group theory, and fluid mechanics. Specifically, after quotienting parameter symmetries and applying local-entropy coarse-graining, the effective dynamics satisfy a viscous Hamilton--Jacobi equation on the quotient manifold. Moreover, under the assumption that the raw parameter dynamics can be summarized by a gradient field on the quotiented space, the gradient of the coarse-grained loss function obeys a Burgers-type equation, and shock formation can be established rigorously. We apply our theory to multilayer perceptrons, convolutional neural networks, Transformers, and mean-field networks, and show that they obey the Hamilton--Jacobi or Burgers-type equations. We conjecture that this framework also yields practical diagnostics for deep learning. In architectures such as Transformers, raw parameter norms are often distorted by symmetry redundancy and may therefore be misleading, whereas symmetry-corrected quotient observables provide a principled basis for monitoring, forecasting, and controlling training-phase transitions.
Taiki Miyagawa