cs.LGAug 31, 2026

Sparse Competition during Training For the Emergence of Specialized Modules

Authors: Baptiste Rossigneux, Karim Haroun

Organizations: Univ. Rennes · Inria, IRISA · University of Paris 8

Abstract

Modularity in deep neural networks has been proposed as a means of improving both interpretability and training by promoting disentangled representations and reducing redundancy. In this work, we study the emergence of modular structure through competition dynamics between groups of neurons during training. We introduce a method that (i) maintains near-baseline accuracy, (ii) induces usage-based modularity by sparsely routing inputs to neuron groups, and (iii) encourages specialization of these modules, such that their activations are correlated with input classes. We evaluate the proposed approach on ImageNet-100 and CIFAR-100 and show that with it, specialized modules emerge without module-level supervision. These modules capture a meaningful high-level structure in the data, with individual modules responding to semantic categories (e.g., dogs or vehicles). We also study the emergence of a hierarchical partition of sub-tasks depending on the number of modules. Our results suggest that competitive dynamics can serve as a simple mechanism for inducing functional modularity in standard architectures.

Explore similar work

Jun 8, 2026cs.LG

Beyond Neural Collapse: Task-Intrinsic Geometry Governs Neural Representations in Modular Arithmetic

While neural collapse (NC) predicts that a KK-class-balanced classifier should organize terminal representations as a (K−1)(K-1)-dimensional simplex equiangular tight frame (ETF), modular addition consistently enters a different regime: networks compress to a two-dimensional cyclic geometry in which both classifier weights and token embeddings lie on circles. We refine the explanation of this phenomenon in three directions. First, we formalize a layerwise non-uniform training mechanism: downstream classifier weights are driven by dense cross-entropy gradients into a rank-2 equiangular configuration before upstream embeddings fully reorganize, and once this classifier plane forms, backpropagated feature gradients constrain embedding motion to the same plane while weight decay suppresses orthogonal components. Second, after this subspace locking, the induced in-plane dynamics admit an entropy-regularized transport interpretation on S1S^1; combined with modular-addition labels, this reduces embedding formation to phase alignment, whose minimizers are single-frequency characters of Z/PZ\mathbb{Z}/P\mathbb{Z} and hence equal-angle points on a circle. Third, we quantify why this solution prevails over NC: a simplex ETF gains only an O(1)O(1) advantage in cross-entropy, whereas the cyclic rank-2 solution enjoys a Θ(K)Θ(K) advantage under Schatten or weight-decay surrogates, yielding a critical threshold λcrit=Θ(1/K)λ_{\mathrm{crit}} = Θ(1/K). Our results explain both why classifier weights move first and why embeddings subsequently align with them, showing that grokking on modular arithmetic is governed not by maximal separation alone but by a task-structured trade-off between separation, symmetry, and complexity.
Hu Tan, Kuo Gai, Shihua Zhang
Aug 25, 2026cs.LG

Revenge of Monosemanticity: Neuron Specialization as a New Form of Feature Learning in MLPs

Understanding how neural networks learn and organize features is central to understanding their behavior. Much existing theory of feature learning has focused on the emergence of a global low-dimensional representation. We show that this picture is incomplete. In regression problems with clustered data, we demonstrate that multilayer perceptrons (MLPs) naturally develop monosemantic specialized neurons: individual neurons become strongly aligned with a specific predictive feature relevant to a particular region of the input space. Rather than learning a single global low-dimensional representation, MLPs learn a collection of local low-dimensional representations. We show that this ability to specialize gives MLPs a provable data-efficiency advantage over feature-learning methods based on a global low-dimensional representation.
Amirhesam Abedsoltan, Enric Boix-Adsera, Fivos Kalogiannis +1
Aug 8, 2026cs.LG

The Neural Division of Labor: Biologically-Inspired Modular Architectures for Robust Neuromorphic Computing

Biological neural systems achieve high efficiency and robustness through compartmentalized architectures. In contrast, modern artificial neural networks rely on globally entangled structures, which obscure decision logic and suffer from catastrophic forgetting. Here, we report a Decomposable Spiking Neural Network (D-SNN) that eliminates global synaptic entanglement by structurally isolating classification pathways into independent experts. Optimized via a bio-inspired push-pull loss function, the D-SNN achieves competitive accuracies on MNIST, Fashion-MNIST, and CIFAR-10/100 benchmarks. This modular approach matches the performance of fully dense networks while utilizing an order of magnitude fewer parameters. In addition, our networks operate with up to several orders of magnitude lower firing rates and fewer synaptic operations. Furthermore, physically severing connections between experts provides inherent protection against catastrophic forgetting during sequential learning. Crucially, these isolated pathways generate auditable neural signals, increasing decision transparency. This biomimetic, verifiable architecture establishes an efficient foundation for deploying deterministic neuromorphic intelligence in resource-constrained edge environments.
Maksim Bazhenov, Serafim Grubas, Vakhtang Putkaradze