Correlational Training of Morphological Neural Networks
Authors: Konstantinos Fotopoulos, Petros Maragos
Organizations: HERON - Hellenic Robotics Center of Excellence, Athena Research Center, Marousi, Greece · Institute of Robotics, Athena Research Center, Marousi, Greece · School of ECE, National Technical University of Athens, Athens, Greece
Neural networks are typically trained using first-order methods and back-propagation. It is unclear whether this approach is optimal for morphological layers whose weight Jacobians are sparse and whose resulting parameter gradients can be poor. In this work, we propose a novel weight update method for morphological neural networks inspired from the Multiplicative Weights Update (MWU) scheme. We view each morphological perceptron as an instance of the learning from experts' advice problem in logarithmic space, and use a correlation-based reward that favors inputs aligned with the desired output change, regardless of whether a strong gradient signal has reached their weight. We empirically evaluate our approach by training fully connected layers both as stand-alone models and as parts of larger transformer networks. Across nine benchmarks, correlational training yields improvements on eight, by up to 32.84 percentage points, while substantially reducing run-to-run variability.
Figures & tables
θ′=θ+ηgθ.
Algorithm 1 MWU-adjusted morphological training.
Dataset
MLP / ViT
Morph. + BP
Morph. + Corr. (ours)
Δ vs. BP (pp)
CIFAR-10
81.08±0.33
63.96±0.88
70.20±0.62
+6.24
Fashion-MNIST
88.94±0.34
62.66±8.18
86.03±0.36
+23.37
MNIST
98.19±0.20
62.46±8.90
95.30±0.38
+32.84
Adult
78.99±1.15
59.14±12.52
78.13±1.08
+18.99
Chess
98.96±0.33
68.70±8.29
87.89±12.69
+19.19
Ionosphere
90.48±2.94
78.38±16.03
83.65±4.74
+5.27
Table 1: Test performance of linear models, morphological models trained using ordinary back-propagation (BP), and the proposed correlational update, both accelerated using Adam, across several benchmarks. The fourth column corresponds to percentage-point difference with respect to ordinary BP. For image datasets we report accuracy; for PMLB balanced accuracy.
Figure 1: Training curves (mean ± std) of linear models (blue), morphological models under vanilla Adam (orange), and morphological models under our proposed method (green). Our method improves training dynamics over vanilla Adam.
The linear recurrent neural network (LRNN) is a simple model for studying how much memory a network builds up as it trains. For uncorrelated inputs, earlier work found that training itself settles the network between keeping the past and reacting only to the present. Real sequences are correlated, and we solve the learning dynamics exactly for correlated inputs. In the solution, keeping the past carries a cost. The whole effect of correlation lands on that cost. This cost reduces to the earlier one when inputs are uncorrelated and grows once they are positively correlated. Three findings follow. (1) Correlation reshapes the course of learning, not only its end. Memory builds, overshoots, and is partly removed, and the settled network keeps less of the past. (2) Memory switches off at a threshold set by one number, how much each input resembles the one just before it. Neither sequence length nor longer-range correlation moves this threshold. Memory is worth keeping only when the task needs the previous input more than the current input already supplies it through correlation with the past. (3) The best network changes too. Zero error demands a feedthrough, a path that passes the current input straight to the network's output and remembers nothing, and training builds it unprompted when given one spare hidden dimension. Our work turns one property of the input into a prediction of whether a network learns memory and explains why correlated data turns recurrent networks into change detectors.
Arnol Manuel Fokam, Fasseu Sieyondji Akpevwoghene, Edem Fiifi Dawson
Independent Researcher, United Kingdom · Department of Electrical and Electronic Engineering, University of Buea, Cameroon · minoHealth AI Labs, Ghana
Recent advances in deep reinforcement learning (RL) have shown that improving neural network architectures can yield substantial gains in sample efficiency and asymptotic performance without altering the underlying algorithms. In contrast, work on multi-objective reinforcement learning (MORL), which aims to discover a set of policies that balance trade-offs among conflicting objectives, has predominantly focused on algorithmic innovations, leaving the area of architectures underexplored. While the optimal policies and value functions can differ significantly depending on the trade-offs, MORL algorithms commonly represent them with simple feedforward networks conditioned on the trade-off. This raises the question of whether the performance of the algorithms could be improved with more expressive function approximators. In this paper, we integrate recent advances in neural network design: (i) observation and feature normalization, (ii) weight normalization, and (iii) modeling of distributional returns with an entropy-regularized MORL algorithm. The empirical results across standard continuous control benchmarks demonstrate that these changes substantially improve the quality of the produced solution sets without requiring major changes to the underlying algorithm.
Adam Štafa, Santeri Heiskanen, Petr Novotný +1
Faculty of Informatics, Masaryk University · Department of Electrical Engineering and Automation, Aalto University
Learning from imperfect data is a central theme in machine learning, connecting practical questions of robustness to fundamental questions of learnability. Here we examine attribute noise: learning from corrupted inputs while keeping the labels intact, a setting that has received considerably less analytical attention than its label-noise counterpart. We consider two types of corruption models: additive noise and replacement noise. Through experiments with multi-layer perceptrons (MLPs) on corrupted classification datasets, we find that neural networks remain robust, maintaining well-above-chance accuracy even when inputs are >90% corrupted -- far beyond human recognition. To understand this robustness, we analyze infinite-width networks in the heavy-corruption regime using a mean-field-inspired approach and derive a leading-order decision rule for the classification outcome: the network implements a prototype rule, the nearest-class-mean, assigning each test point to the class whose training-set average it most closely resembles. This leading-order decision rule is universal across a broad range of MLP architectures, holding for any depth, as well as a wide class of activation functions and noise distributions. The same centroid mechanism closely matches finite-width network behavior in our experiments and provides an interpretable and analytically tractable account of why learning can succeed even when individual training examples carry almost no signal.
Justin Tahmassebpur, Asadullah Bhuiyan, Hyejin Kim +1
Department of Applied and Engineering Physics, Cornell University, Ithaca, NY 14853, USA · Department of Physics, Cornell University, Ithaca, NY 14853, USA