Prototype-Rule Neurosymbolic Regularization for Rank-Constrained Tensor Neural Networks under Label Scarcity
Authors: Eftychios Protopapadakis, Konstantinos Makantasis, Konstantinos M. Giannoutakis
Organizations: Department of Applied Informatics, University of Macedonia, 156 Egnatia Street, 54636 Thessaloniki, Greece · Department of Artificial Intelligence, Faculty of Information & Communication Technology, University of Malta, Msida, Malta
Rank-constrained tensor neural networks reduce the parameterization of high-order inputs, but they do not explicitly constrain class geometry in the learned representation. This study investigates whether a differentiable prototype-rule can provide a complementary inductive bias for Rank-R tensor learning under limited supervision. The proposed framework augments the Rank-R objective with prototype-based regularization and optionally fuses prototype evidence with neural logits at inference. Four hyperspectral benchmarks are evaluated with four Rank-R configurations under both seven-fold stratification and spatially separated folds that mitigate leakage; a separate spatial study varies the class support budget from 2 to 20 samples. Under spatial evaluation, full neurosymbolic inference changes Macro-F1 score by +8.82 percentage points on Botswana, +5.49 on Indian Pines, +1.59 on Pavia University, and -0.62 on Salinas. Most of the benefit arises from training-time regularization, whereas inference fusion is small and dataset dependent.
Figures & tables
Approach
Tensor structure
Explicit relation
Training constraint
Prototype inference
Rank-R FNN
✓
–
–
–
Prototype-based learning
varies
–
varies
varies
Differentiable constraint-based NS
varies
✓
✓
varies
Present framework
✓
✓
✓
optional
Table 1: Conceptual positioning of the proposed mechanism. The table compares mechanisms rather than predictive performance.
Figure 1: High-level overview of the proposed neurosymbolic Rank-R framework. Hyperspectral input patches are mapped to a latent representation through the CP-constrained Rank-R backbone. The learned embedding can be further shaped by the prototype-rule constraint, which introduces class-prototype attraction and inter-class separation during training. The resulting representation supports three matched variants: the purely neural RankR baseline, RankR-NS-RegOnly using prototype-rule regularization during training, and RankR-NS additionally incorporating prototype evidence during inference.
Method
CE
Rule regularization
Prototype fusion
Scientific role
RankR
✓
–
–
Purely neural Rank-R baseline
RankR-NS-RegOnly
✓
✓
–
Training-time rule effect
RankR-NS
✓
✓
✓
Regularization plus inference fusion
Table 2: Experimental variants and mechanistic interpretation.
Dataset
Spatial size
Bands
Classes
Labeled pixels
Patch
Pavia University
610×340
103
9
42,776
5×5
Indian Pines
145×145
200
16
10,249
5×5
Salinas
512×217
204
16
54,129
5×5
Botswana
1476×256
145
14
3,248
5×5
Table 3: Hyperspectral datasets used in the experiments. “Bands” denotes the usable channels in the corrected data used by the experiments.
Figure 2: Spatial split configurations used for leakage-controlled evaluation: (a) Botswana, (b) Indian Pines, (c) Pavia University, and (d) Salinas. Blue and green indicate the selected training and validation samples, respectively, drawn from spatially separated eligible regions. Red denotes all valid labeled samples within the spatially held-out test block. Yellow indicates the spatial dead zone introduced to prevent patch overlap between split roles. Light-gray pixels correspond to labeled samples not selected for training, validation, or testing in the illustrated fold, whereas white denotes unlabeled background. Owing to the highly elongated geometry of the Botswana scene, panel (a) shows an enlarged central crop for visual clarity; the split itself is constructed on the complete scene.
Protocol
Dataset
RankR
RankR-NS-RegOnly
RankR-NS
Spatial
Pavia University
62.57±15.28
64.20±18.15
64.16±17.96
Indian Pines
24.91±2.26
30.19±0.72
30.40±0.02
Salinas
58.04±9.24
58.18±9.15
57.42±8.36
Botswana
48.64±5.21
56.58±1.02
57.46±2.88
Stratified
Pavia University
68.04±1.36
69.41±1.50
69.23±1.61
Indian Pines
51.77±1.68
52.31±1.08
51.82±0.90
Table 4: Absolute dataset-level Macro-F1 (mean ± SD, percentage points). For each method, architecture-specific scores are first averaged within each matched fold and the resulting fold-level values are then summarized across folds.
Protocol
Dataset
RankR-NS-RegOnly − RankR
RankR-NS − RankR
RankR-NS − RankR-NS-RegOnly
Spatial
Pavia University
+1.63
+1.59
-0.05
Indian Pines
+5.28
+5.49
+0.21
Salinas
+0.13
-0.62
-0.76
Botswana
+7.94
+8.82
+0.88
Stratified
Pavia University
+1.36
+1.18
-0.18 ∗
Indian Pines
+0.54
+0.05
-0.49
Table 5: Dataset-level paired change in Macro-F1 (percentage points). Architecture scores are averaged within matched folds before contrasts are formed. Positive values favor the method named first. An asterisk (*) denotes a Holm-corrected post-hoc p<0.05 in the seven-fold stratified analysis. No inferential tests are reported for the two/three-fold spatial protocol.
Figure 3: Paired dataset-level Macro-F1 changes under the spatially separated protocol. The small number of spatial folds is treated descriptively; the figure emphasizes effect direction and magnitude rather than fold-level significance.
Figure 4: Paired dataset-level Macro-F1 changes under seven-fold stratified evaluation. Post-hoc testing is performed only when the dataset-level Friedman omnibus test is significant.
Figure 5: Architecture-level directional Macro-F1 effects for spatial (left) and stratified (right) evaluation. The heatmaps show that architecture averaging can conceal changes in both magnitude and direction.
K
2
3
5
7
10
15
20
RankR-NS − RankR
0.53
0.56
0.97
0.21
0.61
0.46
0.49
RankR-NS-RegOnly − RankR
0.31
0.43
0.62
0.24
0.38
0.39
0.42
Table 6: Equal-dataset-weighted paired Macro-F1 gain over RankR across the label budgets. Values are percentage points and are averaged over the four Rank-R architectures after dataset-level aggregation.
Figure 6: Equal-dataset-weighted full-neurosymbolic Macro-F1 gain over RankR across label budgets and Rank-R architectures. The overall gain is usually positive but not monotonic in K .
Figure 7: Dataset-specific RankR-NS gain over RankR across label budgets. The principal result is heterogeneity rather than a universal scarcity curve.
Tensor-valued data arise naturally in neuroimaging, genomics, climate science, and spatiotemporal networks, where multilinear dependencies across modes carry information that is destroyed under vectorization. Existing approaches either impose a single low-rank structure, which can miss localized signal, or treat the tensor as a long vector, which discards its multiway geometry. We propose a Dual-Channel Tensor Neural Network (DC-TNN) that decomposes each tensor input into a low-rank core and a sparse refinement, and processes the two components through coupled neural channels. The framework is structure-agnostic and accommodates CP, Tucker, and tensor-train cores within a single architecture. For estimation, we establish non-asymptotic risk bounds for the DC-TNN estimator that decompose into network approximation, core estimation, and refinement-selection terms, and show that the effective dimension is determined jointly by the core rank and refinement sparsity rather than by the ambient tensor size. For inference, we develop a structure-aware conformal ROC procedure that calibrates within the core-refinement latent space and produces ROC and AUC confidence bands with finite-sample, distribution-free coverage. Building on this, we propose a conformal structure selector that, to our knowledge, is the first distribution-free procedure for choosing among candidate tensor decompositions with finite-sample validity. Simulations and an analysis of a protein dataset demonstrate competitive predictive accuracy, reliable uncertainty quantification, and consistent recovery of the tensor structure.
Elynn Chen, Jiayu Li, Zheshi Zheng +1
New York University · University of Michigan · Duke University
This paper addresses low-rank tensor completion (LRTC) by proposing a novel nonconvex surrogate, namely the ratio of the tensor nuclear norm to the tensor Ky Fan p-k norm (TNPK), to accurately approximate the tensor tubal rank. The TNPK possesses appealing properties, including scale invariance, parameter flexibility, and the existence of closed-form solutions under specific choices of p and k. With specific parameter settings of p and k, it reduces to the ratio of the tensor nuclear norm to the tensor Ky Fan k norm (TNK) or the ratio of the tensor nuclear norm to the tensor Frobenius norm (TNF). We construct a LRTC model and, under the tensor null space property (NSP), prove that low-rank tensors are local minimizers of the proposed model. Moreover, we derive the proximal operator of the Ky Fan p-k inverse-norm and further develop an efficient alternating direction method of multipliers (ADMM) algorithm with guaranteed subsequential convergence under mild conditions. Extensive experiments on synthetic and real-world datasets validate the superior performance of our method against state-of-the-art competitors.
Shan Fan, Feng Zhang, Jianjun Wang +2
School of Mathematics and Statistics, Southwest University, Chongqing 400715, China · School of Mathematical Sciences/Research Center for Image and Vision Computing, University of Electronic Science and Technology of China, Chengdu 611731, China · Faculty of Computer Science and Control Engineering, Shenzhen University of Advanced Technology, Shenzhen 518055, China
Tensor networks provide efficient representations for compressing large neural networks. By carefully designing shapes and topologies, they can significantly reduce memory and computational costs. However, identifying implicit low-rank structures in large foundation models remains challenging due to their enormous scale and un-structured weight distributions. We propose an adaptive tensorization method that discovers inherent low-rank structure in a target tensor by index ordering. Experiments on weight and KV-cache compression demonstrate improved reconstruction quality compared to baselines.
Toshiaki Koike-Akino, Jing Liu, Ye Wang
Mitsubishi Electric Research Laboratories (MERL), 20 Broadway, Cambridge, MA 02139, USA.