KANFIS: A Neuro-Symbolic Framework for Interpretable and Uncertainty-Aware Learning
Authors: Binbin Yong, Haoran Pei, Jun Shen, Haoran Li, Qingguo Zhou, Zhao Su
Organizations: School of Information Science and Engineering, Lanzhou University, China · School of Computing and Information Technology, University of Wollongong, Australia · Department of Data Science and Artificial Intelligence, Monash University, Australia
Adaptive Neuro-Fuzzy Inference System (ANFIS) was designed to combine the learning capabilities of neural network with the reasoning transparency of fuzzy logic. However, conventional ANFIS architectures suffer from structural complexity, where the product-based inference mechanism causes an exponential explosion of rules in high-dimensional spaces. We herein propose the Kolmogorov-Arnold Neuro-Fuzzy Inference System (KANFIS), a compact neuro-symbolic architecture that unifies fuzzy reasoning with additive function decomposition. KANFIS employs an additive aggregation mechanism, under which both model parameters and rule complexity scale linearly with input dimensionality rather than exponentially. Furthermore, KANFIS is compatible with both Type-1 (T1) and Interval Type-2 (IT2) fuzzy logic systems, enabling explicit modeling of uncertainty and ambiguity in fuzzy representations. By using sparse masking mechanisms, KANFIS generates compact and structured rule sets, resulting in an intrinsically interpretable model with clear rule semantics and transparent inference processes. Empirical results demonstrate that KANFIS achieves competitive performance against representative neural and neuro-fuzzy baselines.
Figures & tables
Figure 1 : Limitations of Existing Interpretable Models. Conventional ANFIS and KAN models suffer from several limitations in terms of interpretability and model complexity.
Figure 2 : Top: KANFIS structures with one and two layers, respectively. Dashed lines indicate paths that may be selected during learning. Bottom: Computational flow of the model. x denotes the input features, y the predicted output, and ω the weight of each rule, reflecting the influence of each rule pattern on the final prediction.
Dataset
Metrics
IT2-KANFIS
T1-KANFIS
MLP
T1-ANFIS
IT2-ANFIS
KAN
CCPP
MAPE
0.7047
0.6777
0.7164
0.6759
0.6696
0.6389
RMSE
4.1240
3.9542
4.1883
3.9980
4.0047
3.9313
MAE
3.1975
3.0760
3.2553
3.0688
3.0439
2.8988
Parkinsons
MAPE
14.891
14.477
19.604
16.167
15.446
14.610
RMSE
0.0397
0.0405
0.0527
0.0449
0.0433
0.0404
MAE
0.0291
0.0289
0.0389
0.0320
0.0310
0.0293
Table 1 : Comparative experimental results. We adopt a clustering-based IT2-ANFIS to ensure its suitability as a baseline under high-dimensional feature settings. The symbol ‘-’ indicates that the model cannot be trained due to the curse of dimensionality.
IF
THEN
Interpretation
AT is HIGH
−0.3407⋅(7.26⋅MAT(xAT))
High ambient temperature reduces air density and mass flow, lowering power output ( Saravanamuttoo et al., 2001 ) .
Low exhaust vacuum and high relative humidity reduce back pressure and increase specific heat, creating optimal operating conditions ( Cengel and Boles, 2002 ) .
AT is MED & AP is HIGH
0.2541⋅(3.38⋅MAT(xAT)+0.48⋅MAP(xAP))
Higher ambient pressure increases air density, enhancing turbine intake flow and efficiency even at moderate temperatures ( Saravanamuttoo et al., 2001 ) .
V is LOW & RH is LOW
0.2038⋅(2.31⋅MV(xV)+1.61⋅MRH(xRH))
Low exhaust vacuum dominates efficiency; reduced condenser back pressure ensures high power output even at low humidity ( Moran et al., 2010 ) .
AT is MED & V is HIGH
−0.1831⋅(3⋅MAT(xAT)+0.53⋅MV(xV))
High exhaust vacuum increases condenser back pressure, restricting steam expansion and offsetting medium-temperature benefits ( Moran et al., 2010 ) .
AT is HIGH & RH is LOW
−0.1351⋅(3.88⋅MAT(xAT)+0.93⋅MRH(xRH))
Low relative humidity implies lower specific heat, slightly reducing gas turbine output at similar temperatures ( Cengel and Boles, 2002 ) .
Table 2 : Interpretability analysis of fuzzy rules on the CCPP dataset. The six most important rules learned by the model and their associated physical principles are presented in the table, with their validity supported by evidence from the cited relevant physical studies.
Figure 3 : Comparison of the average number of features per rule across datasets. The figure illustrates the comparison between the number of features used in the rules and the original number of features in each dataset, depending on whether regularization is applied.
High SBP with low glucose and normal/low temp indicates moderate risk ( D’Agostino et al., 2008 ) .
Age is HIGH
Phighrisk=−0.3013⋅(2.81⋅MAge(xAge))
Advanced age independently increases cardiovascular risk ( D’Agostino et al., 2008 ) .
Table 3 : Interpretability analysis of fuzzy rules on the MHR dataset. The table presents the six most important rules learned by the model along with their corresponding medical principles, and their validity is supported by evidence from the cited relevant studies. For clarity, only the class formula that best matches each rule is reported in the THEN clause.
Dataset
Metric
Normal Regularization
High Regularization
Low Regularization
CCPP
RMSE
4.1240
4.2304
4.1903
Parkinsons
RMSE
0.0397
0.0418
0.0403
BCW
Acc
0.9912
0.9474
0.9474
Table 4: Impact of Different Levels of Regularization on Model Accuracy.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Dimension
T1-KANFIS (Acc)
IT2-KANFIS
APS Failure at Scania Trucks
171
0.9898
0.9900
Dexter
20000
0.8667
0.9333
ForestCover
54
0.9420
0.9491
Appendix
Table 5: Accuracy of KANFIS on the supplementary datasets. The three additional datasets reported in the table are all highly high-dimensional datasets. Among them, the ForestCover dataset was specifically included due to its extremely large scale, containing 581,012 instances.
Dataset
Metric
Ours
NAMs
GAMI-NET
NODE-GAM
Parkinsons
RMSE
0.0397
0.0408
0.0403
–
BCW
Acc
0.9912
0.9651
0.9649
0.9737
Appendix
Table 6: Performance comparison with other interpretable models. The table presents comparative results with three representative interpretable models, namely NAMs, GAMI-NET, and NODE-GAM. Due to limitations in the available open-source implementation, NODE-GAM was unable to produce results on the Parkinsons dataset.
Dataset
IT2-KANFIS
T1-KANFIS
MLP
ANFIS
KAN
CCPP
0.3914, 4.1240
0.3560, 3.9542
0.0902, 4.1883
0.3042, 3.9980
4.8063, 3.9313
BCW
0.0212, 0.9912
0.0207, 0.9912
0.0067, 0.9912
–
1.4943, 0.9736
Spam
0.1357, 0.9392
0.1175, 0.9381
0.0475, 0.9175
–
–
Appendix
Table 7: Time-Accuracy Comparison of Models. The training time is measured in seconds per training epoch. For regression datasets, accuracy is measured using RMSE, while for classification datasets, accuracy is measured using Acc. Values are reported as RMSE and Accuracy.
Dataset
IT2-KANFIS
T1-KANFIS
MLP
ANFIS
KAN
CCPP
521, 2600
361, 1640
193, 480
429, 1274
1600, 9972
BCW
4822, 24330
3986, 18396
1058, 2208
5139, 15398
2240, 11518
Spam
7314, 36960
5034, 23280
1922, 3936
–
–
Appendix
Table 8: Comparison of Parameter Count and FLOPs for Different Models. Values are reported as the number of parameters and FLOPs.
The adaptive neuro-fuzzy inference system (ANFIS) is an interpretable reasoning framework capable of generating explicit IF-THEN fuzzy rules, making it suitable for tasks requiring transparent reasoning. However, existing ANFIS models generally construct rule antecedents and perform inference in Euclidean space, limiting their representational capacity and predictive performance. To address this issue, we propose Hyperbolic ANFIS (HyperANFIS), a hyperbolic extension of ANFIS. HyperANFIS preserves the fuzzy semantics and core architecture of conventional ANFIS while performing rule-prototype learning, rule activation, and consequent aggregation in hyperbolic space. It also retains the ability to generate interpretable IF-THEN rules. By exploiting the representational properties of hyperbolic geometry, HyperANFIS strengthens the fuzzy inference process, thereby improving predictive accuracy, inter-rule collaboration, and the credibility of its interpretable rules. Experimental results show that HyperANFIS consistently outperforms the standard ANFIS baseline and various ANFIS variants across all datasets, while also generating higher-quality fuzzy rules.
Haoran Pei, Zhao Su, Zetao Lin +6
school of Information Science and Engineering, Lanzhou University · Department of Data Science and AI, Monash University · School of Computing and Information Technology, University of Wollongong +1
Kolmogorov--Arnold Networks (KANs) replace scalar edge weights with learnable univariate functions parameterized by multiple basis coefficients. This introduces a source of redundancy that conventional neural-network compression does not directly expose. We present \textbf{SparseKAN}, a unified approach that compresses KANs along three complementary axes: basis functions, neurons/channels, and numerical precision. SparseKAN equips the base branch, nonlinear basis branch, and individual basis terms with hierarchical learnable gates trained under a differentiable active-cost objective. The learned importance structure is subsequently hardened under explicit basis and width budgets, recovered in full or low precision, and physically compacted into smaller dense tensors rather than retained as sparse masks. Experiments on MNIST, CIFAR-10, and CIFAR-100 across spline, polynomial, RBF, wavelet, and convolutional KAN variants show that the structural axes compose predictably in cost. We also find strong basis-dependent differences in term importance: coefficient-based selection outperforms matched low-order truncation by up to 15.25 accuracy points in the evaluated Gram-polynomial settings. Eight-bit quantization is broadly robust, whereas 4-bit convolutional KANs require quantization-aware adaptation. Physical compaction removes up to 73.0% of parameters without accuracy loss on MNIST and reduces large-batch CUDA latency to as little as 0.51× dense execution. On a ZCU104 FPGA, the resulting sparse low-bit models achieve up to 23.63× lower inference latency, demonstrating that SparseKAN converts functional redundancy into measurable software and hardware efficiency. The SparseKAN implementation is available at https://github.com/OSU-STARLAB/SparseKAN.
Kazi Ahmed Asif Fuad, Lizhong Chen
Department of EECS Oregon State University Corvallis, OR 97331
Interpretable machine learning is essential in high-stakes domains where decision-making requires accountability, transparency, and trust. While rule-based models offer global and exact interpretability, learning rule sets that simultaneously achieve high predictive performance and low, human-understandable complexity remains challenging. To address this, we introduce TT-Sparse, a flexible neural building block that leverages differentiable truth tables as nodes to learn sparse, effective connections. A key contribution of our approach is a new soft TopK operator with straight-through estimation for learning discrete, cardinality-constrained feature selection in an end-to-end differentiable manner. Crucially, the forward pass remains sparse, enabling efficient computation and exact symbolic rule extraction. As a result, each node (and the entire model) can be transformed exactly into compact, globally interpretable DNF/CNF Boolean formulas via Quine-McCluskey minimization. Extensive empirical results across 28 datasets spanning binary, multiclass, and regression tasks show that the learned sparse rules exhibit superior predictive performance with lower complexity compared to existing state-of-the-art methods.
Hans Farrell Soegeng, Sarthak Ketanbhai Modi, Thomas Peyrin
School of Physical and Mathematical Sciences, Nanyang Technological University, Singapore.