cs.LGSep 30, 2026

The Conflict Between Logic and Memory: Learning Higher-Order Interactions in Shallow MLPs

Authors: Gongyue Zhang, Honghai Liu

Abstract

A network can fit its training examples while failing to recover the rule that generated their labels. We examine this separation in single-hidden-layer multilayer perceptrons (MLPs), using synthetic tasks that control interaction order and the presence of nuisance inputs. We establish elementary benchmark properties: pure parity contains no predictive lower-order marginals, admits an exact Bayes posterior, and can be represented on clean latent inputs by a width-kk ReLU network. Experiments then identify distinct optimization outcomes. In a matched order-2--4 sweep, SGD, Adam, and Muon all reach 100% peak test accuracy at order two; at order three they reach 96.25%, 50.87%, and 76.82%, respectively, while Muon reaches 99.21% at order four. In a separate mixed-order task, freezing only the first-layer weights connected to independent nuisance inputs raises AdamW's epoch-10 accuracy from 44.73% to 95.07%. Removing the same inputs only at test time raises it to 48.38%. Thus, nuisance-weight learning changes the training outcome beyond its immediate effect on prediction. Bias interventions expose a connection between target symmetry and shallow ReLU representations. In a compact signal-only regime, both SGD and Muon learn orders five through eight, with higher SGD peak accuracy at orders nine through eleven. Together, the results show how optimization and nuisance learning constrain the higher-order rules realized by a shallow network.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Implicit Bias of SGD in Multivariate ReLU Networks: Effective Width Collapse

    Jul 3, 2026Shuang Liang, Tom Jacobs, Guido MontúfarRectified Linear Unit NetworksStochastic Gradient Descent

  2. Revenge of Monosemanticity: Neuron Specialization as a New Form of Feature Learning in MLPs

    Aug 25, 2026Amirhesam Abedsoltan, Enric Boix-Adsera, Fivos Kalogiannis +1Feature LearningMultilayer Perceptrons

  3. Let the Neurons Die: Exploiting ReLU-Induced Model Degradation

    Sep 28, 2026Kexin Li, Wenjun Qiu, Joshua Abraham +2Rectified Linear Unit NetworksNeurons