Correlational Training of Morphological Neural Networks
Authors: Konstantinos Fotopoulos, Petros Maragos
Organizations: HERON - Hellenic Robotics Center of Excellence, Athena Research Center, Marousi, Greece · Institute of Robotics, Athena Research Center, Marousi, Greece · School of ECE, National Technical University of Athens, Athens, Greece
Neural networks are typically trained using first-order methods and back-propagation. It is unclear whether this approach is optimal for morphological layers whose weight Jacobians are sparse and whose resulting parameter gradients can be poor. In this work, we propose a novel weight update method for morphological neural networks inspired from the Multiplicative Weights Update (MWU) scheme. We view each morphological perceptron as an instance of the learning from experts' advice problem in logarithmic space, and use a correlation-based reward that favors inputs aligned with the desired output change, regardless of whether a strong gradient signal has reached their weight. We empirically evaluate our approach by training fully connected layers both as stand-alone models and as parts of larger transformer networks. Across nine benchmarks, correlational training yields improvements on eight, by up to 32.84 percentage points, while substantially reducing run-to-run variability.
Figures & tables
θ′=θ+ηgθ.
Algorithm 1 MWU-adjusted morphological training.
Dataset
MLP / ViT
Morph. + BP
Morph. + Corr. (ours)
Δ vs. BP (pp)
CIFAR-10
81.08±0.33
63.96±0.88
70.20±0.62
+6.24
Fashion-MNIST
88.94±0.34
62.66±8.18
86.03±0.36
+23.37
MNIST
98.19±0.20
62.46±8.90
95.30±0.38
+32.84
Adult
78.99±1.15
59.14±12.52
78.13±1.08
+18.99
Chess
98.96±0.33
68.70±8.29
87.89±12.69
+19.19
Ionosphere
90.48±2.94
78.38±16.03
83.65±4.74
+5.27
Table 1: Test performance of linear models, morphological models trained using ordinary back-propagation (BP), and the proposed correlational update, both accelerated using Adam, across several benchmarks. The fourth column corresponds to percentage-point difference with respect to ordinary BP. For image datasets we report accuracy; for PMLB balanced accuracy.
Figure 1: Training curves (mean ± std) of linear models (blue), morphological models under vanilla Adam (orange), and morphological models under our proposed method (green). Our method improves training dynamics over vanilla Adam.
Department of Applied and Engineering Physics, Cornell University, Ithaca, NY 14853, USA · Department of Physics, Cornell University, Ithaca, NY 14853, USA