Parameter space symmetries and conservation laws play an important role in understanding the loss landscapes and implicit biases of neural networks. Inspired by Noether's theorem in physics, prior works have sought to derive conservation laws under gradient flow from parameter symmetries, but the scope and limitations of this connection remain unclear. We develop a unified geometric framework that clarifies the precise relationship between the two notions, including the conditions under which symmetries correspond to conservation laws. We introduce a notion of compositional identifiability and use it to establish a general inheritance principle for complete characterizations of symmetries and conservation laws in multilayer networks. We apply the framework to multi-head and grouped-query attention, polynomial neural networks, and square deep linear networks.
Figures & tables
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Combination Type
Operation
Typical Examples
Additive
z=gθ(x)+fω(x)
ResNet [ He et al., 2015 ] , LoRA [ Hu et al., 2022 ]
Concatenation
z=[gθ(x);fω(x,gθ(x))]
Inception [ Szegedy et al., 2015 ] , U-Net [ Ronneberger et al., 2015 ]
Composition
z=fω(gθ(x)), or z=gθ(fω(x))
MLPs [ Rumelhart et al., 1986 ] , CNNs [ Lecun et al., 1998 ]
Hadamard Product / Gating
z=gθ(x)⊙fω(x)
LSTM [ Hochreiter and Schmidhuber, 1997 ] and GRU gates [ Cho et al., 2014 ]
Input-dependent Weighted Sum
z=aω(x)gθ(x)+fω(x)
Mixture-of-Experts [ Shazeer et al., 2017 ]
Outer Product
Z=gθ(x)fω(x)⊤
Bilinear CNNs [ Lin et al., 2015 ]
Appendix
Table 1: Examples of function combinations compatible with Proposition 15 . The listed operations occur within the indicated architectures. The categories are not mutually exclusive. The list is not exhaustive. We write gθ(u)=g(θ,u) , and N(i) is the neighborhood of node i .
We explore whether intrinsic symmetries of the training data lead to conserved quantities during gradient-flow training of neural networks. Under the assumption that the loss function is analytic and non-polynomial, we prove that data symmetries generically do not induce any additional integrals of motion. For mean squared error (MSE) loss, on the other hand, there are situations in which data augmentation yields extra conserved quantities. We build a framework, utilizing \emph{tensorizable networks} to describe this phenomenon. Tensorizable networks are a family of architectures whose dependence on parameters and inputs can be separated using an intermediate representation. They include linear and polynomial networks, as well as Lightning Attention.
Understanding gradient descent dynamics is key to explaining the success of over-parameterized models, where implicit bias manifests through conservation laws in gradient flow. While such laws are well understood for linear and ReLU networks, they remain largely unexplored for modern architectures. This work develops a unified framework to characterize conservation laws for contemporary models, including feedforward networks with GELU, SiLU, and SwiGLU activations, multihead attention with sinusoidal and rotary positional encodings, and Mixture-of-Experts architectures under diverse gating designs. Our theoretical findings are supported by experiments that validate the predicted invariants.
Viet-Hoang Tran, Vinh Khanh Bui, Tan Lai Ngoc +3
National University of Singapore · Center for AI Research, VinUniversity · Independent Researcher +1
Parameter space symmetries are important for understanding neural networks' loss landscape, training dynamics, and generalization. However, systematically identifying these symmetries remains a challenge. In this paper, we formalize data-dependent parameter symmetries and characterize loss invariance and the group-action axioms through infinitesimal conditions, which provide objectives for jointly learning group generators and nonlinear action maps. Our framework systematically uncovers parameter symmetries, including previously unknown ones. To study larger networks, we establish conditions under which subnetwork symmetries extend to the full model. The same construction gives an explicit family of finite-batch symmetries, providing both analytical examples and a foundation for discovery through small subnetworks. Using the infinitesimal characterization and subnetwork construction, we implement a framework for automated discovery of parameter symmetries, and successfully uncovered symmetries in various architectures, including pretrained transformer models.
Bo Zhao, Nima Dehmamy, Robin Walters +1
Harvard University · IBM Research · Northeastern University +1