In many high-stakes applications, machine learning is dominated by black-box models that require post hoc explanations to justify their predictions. These explanations are often unreliable because they do not reflect the model's actual computations, limiting accountability and trust. A natural alternative is to use models that are interpretable by design. However, existing rule-based approaches, such as RuleFit and decision trees, while transparent, often lack stability and predictive strength, reinforcing a perceived trade-off between traditional performance measures and model understandability. To address this, we propose Sparse Relaxed Regularized Regression Rule-Fit (SR4-Fit), an intrinsically interpretable algorithm for both classification and regression that produces compact and stable rule sets without sacrificing performance. Using demographic data from the U.S. Census Bureau's American Community Survey, SR4-Fit predicts U.S. House election outcomes with high accuracy and interpretability while uncovering demographic interactions missed by black-box models. We further validate SR4-Fit across fourteen benchmark datasets (six classification and eight regression), where it outperforms existing rule-based methods, including RuleFit and decision trees in terms of accuracy, stability, and compactness while remaining competitive with black-box models in predictivity. These results demonstrate that interpretability and predictive reliability need not be mutually exclusive, offering a practical and transparent alternative for high-stakes decision-making.
Figures & tables
Figure 1 : Conceptual illustration of the predictivity–understandability trade-off across common machine learning models. SR4-Fit is positioned closer to the ideal region of high predictivity and high understandability, reflecting its aim to balance both criteria more effectively than existing rule-based approaches.
Figure 2 : Overview of the SR4-Fit pipeline (including classification and regression).
β=argminβ21∥y−Zβ∥22+2κ∥β−w∥22
Algorithm 2 SR4-Fit Regression Algorithm
Figure 3 : Aggregate Pareto plots for the U.S. House election classification and regression tasks, showing mean predictivity (F1 score for classification, left; R2 score for regression, right) against Rule Understandability Score (RUS). The classification panel averages results across all four demographic datasets and both feature configurations (with and without party percentages). The regression panel averages results across all four datasets and all four experimental configurations (DEM as label with and without REP percentage; REP as label with and without DEM percentage).
Figure 4 : Aggregate Pareto plots for the public benchmark classification and regression tasks, showing mean predictivity (F1 score for classification, left; R2 score for regression, right) against Rule Understandability Score (RUS). The classification panel averages results across the six public classification benchmark datasets. The regression panel averages results across the eight public regression benchmark datasets.
Table 1 : Average enhanced interpretability score (E-IPS) ± standard deviation across 30 trials for demographic classification datasets. Bold indicates the best-performing method per dataset.
Figure 5 : Distribution of enhanced interpretability scores (E-IPS) across demographic classification datasets. The left panel reports results obtained using feature sets that include party percentage information, whereas the right panel shows results with party percentages removed.
Table 2 : Average enhanced interpretability score (E-IPS) ± standard deviation across 30 trials for demographic regression datasets. Results are reported for predicting Democratic vote percentage, with and without Republican percentage included as a feature. Bold indicates the best-performing method per dataset.
Figure 6 : Distribution of enhanced interpretability scores (E-IPS) across demographic regression datasets when predicting Democratic vote percentage. Results are shown for models trained with all party-related features (left) and for models trained with Republican percentage features removed (right).
Table 3 : Average enhanced interpretability score (E-IPS) ± standard deviation across 30 trials for demographic regression datasets. Results are reported for predicting Republican vote percentage, with and without Democratic percentage included as a feature. Bold indicates the best-performing method per dataset.
Figure 7 : Distribution of enhanced interpretability scores (E-IPS) across demographic regression datasets when predicting Republican vote percentage. Results are shown for models trained with all party-related features (left) and for models trained with Democratic percentage features removed (right).
Dataset
Random Forest
RuleFit
SR4-Fit
Decision Tree
Breast Cancer
0.4350 ± 0.1433
0.4793 ± 0.1851
0.4931 ± 0.1875
0.4477 ± 0.2044
E. coli
0.4505 ± 0.1475
0.6085 ± 0.1087
0.6141 ± 0.1093
0.3624 ± 0.1494
Page Blocks
0.5386 ± 0.1296
0.4813 ± 0.1322
0.5308 ± 0.1381
0.5405 ± 0.1391
Pima Indians
0.4644 ± 0.1330
0.3596 ± 0.3061
0.3208 ± 0.3011
0.3862 ± 0.1781
Vehicle
0.4958 ± 0.1642
0.4175 ± 0.1314
0.4099 ± 0.1262
0.5338 ± 0.1570
Yeast
0.4927 ± 0.1426
0.6313 ± 0.1234
0.6512 ± 0.1281
0.4017 ± 0.1596
Table 4 : Average enhanced interpretability score (E-IPS) ± standard deviation across 30 trials for public classification datasets. Bold indicates the highest average E-IPS per dataset.
Figure 8 : Violin plot comparison of enhanced interpretability scores for all rule-based models across standard public classification datasets.
Dataset
Random Forest
RuleFit
SR4-Fit
Decision Tree
Abalone
0.4890 ± 0.1166
0.4824 ± 0.0649
0.1444 ± 0.1525
0.5524 ± 0.1113
Bone
0.4723 ± 0.1565
0.5756 ± 0.0549
0.3263 ± 0.3564
0.6124 ± 0.2261
Diabetes
0.4019 ± 0.1627
0.6474 ± 0.0588
0.6176 ± 0.2149
0.5338 ± 0.0878
Housing
0.4965 ± 0.1320
0.5330 ± 0.0613
0.5425 ± 0.1730
0.4301 ± 0.1264
Machine
0.5615 ± 0.1073
0.5410 ± 0.1868
0.5588 ± 0.1688
0.4165 ± 0.1849
MPG
0.4738 ± 0.1402
0.5391 ± 0.0615
0.5393 ± 0.2905
0.5655 ± 0.0507
Table 5 : Average enhanced interpretability score (E-IPS) ± standard deviation across 30 trials for public regression datasets. Bold indicates the highest average E-IPS per dataset.
Figure 9 : Violin plot comparison of enhanced interpretability scores for all rule-based models across standard public regression datasets.
Appendix figures & tables72 assets
Supplementary material from the paper’s appendix.
Appendix
Table 6 : Wilcoxon test W-statistic and p-value comparison between RuleFit and SR4-Fit over 30 trials across demographic classification datasets. Statistically significant differences ( p<0.05 ) are shown in bold .
Table 7 : Wilcoxon signed-rank test W-statistic and p-value comparison between RuleFit and SR4-Fit over 30 trials across demographic datasets for Democratic percentage label regression. Statistically significant differences ( p<0.05 ) are shown in bold .
Table 8 : Wilcoxon signed-rank test W-statistic and p-value comparison between RuleFit and SR4-Fit over 30 trials across demographic datasets for Republican percentage label regression. Statistically significant differences ( p<0.05 ) are shown in bold .
Figure 10 : Line plot comparison of model performance metrics for minimum data across 30 trials. The left panel reports results obtained using feature sets that include prior party voting percentage information, whereas the right panel shows results with prior party voting percentages removed.
Figure 11 : Line plot comparison of model (logistic regression, Decision Tree, XG boost) performance metrics for minimum data across 30 trials. The left panel reports results obtained using feature sets that include prior party voting percentage information, whereas the right panel shows results with prior party voting percentages removed.
Figure 12 : Line plot comparison of model performance metrics for standard data across 30 trials. The left panel reports results obtained using feature sets that include prior party voting percentage information, whereas the right panel shows results with prior party voting percentages removed.
Figure 13 : Line plot comparison of model (logistic regression, Decision Tree, XG boost) performance metrics for standard data across 30 trials. The left panel reports results obtained using feature sets that include prior party voting percentage information, whereas the right panel shows results with prior party voting percentages removed.
Figure 14 : Line plot comparison of model performance metrics for expanded data across 30 trials. The left panel reports results obtained using feature sets that include prior party voting percentage information, whereas the right panel shows results with prior party voting percentages removed.
Figure 15 : Line plot comparison of model (logistic regression, Decision Tree, XG boost) performance metrics for expanded data across 30 trials. The left panel reports results obtained using feature sets that include prior party voting percentage information, whereas the right panel shows results with prior party voting percentages removed.
Figure 16 : Line plot comparison of model performance metrics for previous data across 30 trials. The left panel reports results obtained using feature sets that include prior party voting percentage information, whereas the right panel shows results with prior party voting percentages removed.
Figure 17 : Line plot comparison of model (logistic regression, Decision Tree, XG boost) performance metrics for previous data across 30 trials. The left panel reports results obtained using feature sets that include prior party voting percentage information, whereas the right panel shows results with prior party voting percentages removed.
Figure 18 : Violin plot comparison of model performance metrics across multiple datasets for the election classification task that includes prior party voting percentage information.
Figure 19 : Violin plot comparison of model (logistic regression, Decision Tree, XG boost) performance metrics across multiple datasets for the election classification task that includes prior party voting percentage information.
Figure 20 : Violin plot comparison of model performance metrics across multiple datasets for the election classification task that does not include prior party voting percentage information.
Figure 21 : Violin plot comparison of model (logistic regression, Decision Tree, XG boost) performance metrics across multiple datasets for the election classification task that does not include prior party voting percentage information.
Table 9 : Average Dice–Sørensen Index ± standard deviation across 30 trials for election classification datasets. Bold indicates the best-performing method per dataset.
Table 10 : Average number of rules ± standard deviation across 30 trials for election classification datasets. Bold indicates the most compact model.
Table 11 : Average rule complexity (number of conditions per rule) ± standard deviation for election classification datasets across 30 trials. Bold indicates less complex rules.
Figure 22 : Line plot comparison of model performance metrics for minimum data across 30 trials with DEM percentage as label. The left panel reports results obtained with party percentage (REP) included, whereas the right panel shows results with party percentage (REP) not included.
Figure 23 : Line plot comparison of model (LASSO regression, Decision Tree, XG boost) performance metrics for minimum data across 30 trials. The left panel reports results obtained with party percentage (REP) included, whereas the right panel shows results with party percentage (REP) not included.
Figure 24 : Line plot comparison of model performance metrics for standard data across 30 trials with DEM percentage as label. The left panel reports results obtained with party percentage (REP) included, whereas the right panel shows results with party percentage (REP) not included.
Figure 25 : Line plot comparison of model (LASSO regression, Decision Tree, XG boost) performance metrics for standard data across 30 trials. The left panel reports results obtained with party percentage (REP) included, whereas the right panel shows results with party percentage (REP) not included.
Figure 26 : Line plot comparison of model performance metrics for expanded data across 30 trials with DEM percentage as label. The left panel reports results obtained with party percentage (REP) included, whereas the right panel shows results with party percentage (REP) not included.
Figure 27 : Line plot comparison of model (LASSO regression, Decision Tree, XG boost) performance metrics for expanded data across 30 trials. The left panel reports results obtained with party percentage (REP) included, whereas the right panel shows results with party percentage (REP) not included.
Figure 28 : Line plot comparison of model performance metrics for previous data across 30 trials with DEM percentage as label. The left panel reports results obtained with party percentage (REP) included, whereas the right panel shows results with party percentage (REP) not included.
Figure 29 : Line plot comparison of model (LASSO regression, Decision Tree, XG boost) performance metrics for previous data across 30 trials. The left panel reports results obtained with party percentage (REP) included, whereas the right panel shows results with party percentage (REP) not included.
Figure 30 : Violin plot comparison of model performance metrics across multiple datasets for the election regression task that has DEM percentage as a label and includes party percentage (REP).
Figure 31 : Violin plot comparison of model (LASSO regression, Decision Tree, XG boost) performance metrics across multiple datasets for the election regression task that has DEM percentage as a label and includes party percentage (REP).
Figure 32 : Violin plot comparison of model performance metrics across multiple datasets for the election regression task that has DEM percentage as a label and does not include party percentage (REP).
Figure 33 : Violin plot comparison of model (LASSO regression, Decision Tree, XG boost) performance metrics across multiple datasets for the election regression task that has DEM percentage as a label and does not include party percentage (REP).
Table 12 : Average Dice–Sørensen Index ± standard deviation across 30 trials for the election regression dataset where DEM percentage is the label.
Table 13 : Average number of rules ± standard deviation across 30 trials for the election regression dataset, where DEM percentage is the label.
Table 14 : Average rule complexity ± standard deviation across 30 trials for the election regression dataset where DEM percentage is the label.
Figure 34 : Line plot comparison of model performance metrics for minimum data across 30 trials with REP percentage as label. The left panel reports results obtained with party percentage (DEM) included, whereas the right panel shows results with party percentage (DEM) not included.
Figure 35 : Line plot comparison of model (LASSO regression, decision tree, XG boost) performance metrics for minimum data across 30 trials. The left panel reports results obtained with party percentage (DEM) included, whereas the right panel shows results with party percentage (DEM) not included.
Figure 36 : Line plot comparison of model performance metrics for standard data across 30 trials with REP percentage as label. The left panel reports results obtained with party percentage (DEM) included, whereas the right panel shows results with party percentage (DEM) not included.
Figure 37 : Line plot comparison of model (LASSO regression, decision tree, XG boost) performance metrics for standard data across 30 trials. The left panel reports results obtained with party percentage (DEM) included, whereas the right panel shows results with party percentage (DEM) not included.
Figure 38 : Line plot comparison of model performance metrics for expanded data across 30 trials with REP percentage as label. The left panel reports results obtained with party percentage (DEM) included, whereas the right panel shows results with party percentage (DEM) not included.
Figure 39 : Line plot comparison of model (LASSO regression, decision tree, XG boost) performance metrics for expanded data across 30 trials. The left panel reports results obtained with party percentage (DEM) included, whereas the right panel shows results with party percentage (DEM) not included.
Figure 40 : Line plot comparison of model performance metrics for previous data across 30 trials with REP percentage as label. The left panel reports results obtained with party percentage (DEM) included, whereas the right panel shows results with party percentage (DEM) not included.
Figure 41 : Line plot comparison of model (LASSO regression, decision tree, XG boost) performance metrics for previous data across 30 trials. The left panel reports results obtained with party percentage (DEM) included, whereas the right panel shows results with party percentage (DEM) not included.
Figure 42 : Violin plot comparison of model performance metrics across multiple datasets for the election regression task that has REP percentage as a label and includes party percentage (DEM).
Figure 43 : Violin plot comparison of model (LASSO regression, decision tree, XG boost) performance metrics across multiple datasets for the election regression task that has REP percentage as a label and includes party percentage (DEM).
Figure 44 : Violin plot comparison of model performance metrics across multiple datasets for the election regression task that has REP percentage as a label and does not include party percentage (DEM).
Figure 45 : Violin plot comparison of model (LASSO regression, decision tree, XG boost) performance metrics across multiple datasets for the election regression task that has REP percentage as a label and does not include party percentage (DEM).
Table 15 : Average Dice–Sørensen Index ± standard deviation across 30 trials for the election regression dataset where REP percentage is the label.
Table 16 : Average number of rules ± standard deviation across 30 trials for the election regression dataset, where REP percentage is the label.
Table 17 : Average rule complexity ± standard deviation across 30 trials for the election regression dataset, where REP percentage is the label.
Figure 46 : Line plot comparison of model performance metrics for breast cancer data across 30 trials. The left panel reports results obtained for models—random forest, SVM, rulefit, and SR4-fit , whereas the right panel shows results with models logistic regression, decision tree, and XGboost.
Figure 47 : Line plot comparison of model performance metrics for E. coli data across 30 trials. The left panel reports results obtained for models—random forest, SVM, rulefit, and SR4-fit , whereas the right panel shows results with models logistic regression, decision tree, and XGboost.
Figure 48 : Line plot comparison of model performance metrics for page blocks data across 30 trials. The left panel reports results obtained for models—random forest, SVM, rulefit, and SR4-fit , whereas the right panel shows results with models logistic regression, decision tree, and XGboost.
Figure 49 : Line plot comparison of model performance metrics for Pima Indians data across 30 trials. The left panel reports results obtained for models—random forest, SVM, rulefit, and SR4-fit , whereas the right panel shows results with models logistic regression, decision tree, and XGboost.
Figure 50 : Line plot comparison of model performance metrics for vehicle data across 30 trials. The left panel reports results obtained for models—random forest, SVM, rulefit, and SR4-fit , whereas the right panel shows results with models logistic regression, decision tree, and XGboost.
Figure 51 : Line plot comparison of model performance metrics for yeast data across 30 trials. The left panel reports results obtained for models—random forest, SVM, rulefit, and SR4-fit , whereas the right panel shows results with models logistic regression, decision tree, and XGboost.
Figure 52 : Violin plot comparison of model (random forest, SVM, RuleFit, and SR4-fit) performance across multiple standard benchmark classification datasets for different prediction metrics.
Figure 53 : Violin plot comparison of model (logistic regression, decision tree, and XGBoost) performance across multiple standard benchmark classification datasets for different prediction metrics.
Dataset
Random Forest
RuleFit
SR4-Fit
Decision Tree
Breast Cancer
0.3307 ± 0.0186
0.4449 ± 0.0199
0.4449 ± 0.0199
0.0558 ± 0.0540
E. coli
0.1160 ± 0.0150
0.1762 ± 0.0140
0.1762 ± 0.0140
0.0414 ± 0.0347
Page Blocks
0.1090 ± 0.0155
0.3232 ± 0.0294
0.3232 ± 0.0294
0.0135 ± 0.0090
Pima Indians
0.0454 ± 0.0078
0.5153 ± 0.0115
0.5153 ± 0.0115
0.0021 ± 0.0032
Vehicle
0.1993 ± 0.0146
0.5231 ± 0.0272
0.5231 ± 0.0272
0.0593 ± 0.0360
Yeast
0.1182 ± 0.0146
0.1807 ± 0.0182
0.1807 ± 0.0182
0.0121 ± 0.0098
Appendix
Table 18 : Average Dice–Sørensen Index ± standard deviation across 30 trials for public classification datasets. Bold indicates the best-performing method per dataset.
Dataset
Random Forest
RuleFit
SR4-Fit
Decision Tree
Breast Cancer
104.43 ± 11.58
21.13 ± 1.11
21.13 ± 1.11
9.10 ± 0.66
E. coli
117.17 ± 13.48
57.00 ± 0.00
57.00 ± 0.00
11.03 ± 0.67
Page Blocks
50.27 ± 8.77
43.10 ± 2.15
43.10 ± 2.15
18.23 ± 0.89
Pima Indians
37.47 ± 8.19
15.67 ± 0.48
15.67 ± 0.48
20.77 ± 1.96
Vehicle
147.37 ± 11.09
40.07 ± 0.78
40.07 ± 0.78
18.57 ± 1.33
Yeast
57.00 ± 8.87
57.00 ± 0.00
57.00 ± 0.00
24.07 ± 1.41
Appendix
Table 19 : Average number of rules ± standard deviation across 30 trials for public classification datasets. Lower values indicate more compact models.
Dataset
Random Forest
RuleFit
SR4-Fit
Decision Tree
Breast Cancer
1.99 ± 0.01
2.97 ± 0.12
2.97 ± 0.12
3.42 ± 0.15
E. coli
1.97 ± 0.01
3.55 ± 0.01
3.55 ± 0.01
3.60 ± 0.08
Page Blocks
1.97 ± 0.01
2.47 ± 0.04
2.47 ± 0.04
4.41 ± 0.05
Pima Indians
1.96 ± 0.01
2.28 ± 0.07
2.28 ± 0.07
4.61 ± 0.08
Vehicle
1.99 ± 0.01
2.05 ± 0.04
2.05 ± 0.04
4.52 ± 0.09
Yeast
1.96 ± 0.02
2.66 ± 0.03
2.66 ± 0.03
4.69 ± 0.06
Appendix
Table 20 : Average rule complexity (number of conditions per rule) ± standard deviation across 30 trials for public classification datasets. Bold indicates less complex rules.
Figure 54 : Line plot comparison of model performance metrics for abalone data across 30 trials. The left panel reports results obtained for models—random forest, SVM, RuleFit, and SR4-fit , whereas the right panel shows results with models LASSO regression, decision tree, and XGBoost.
Figure 55 : Line plot comparison of model performance metrics for bone data across 30 trials. The left panel reports results obtained for models—random forest, SVM, RuleFit, and SR4-fit, whereas the right panel shows results with models LASSO regression, decision tree, and XGBoost.
Figure 56 : Line plot comparison of model performance metrics for diabetes data across 30 trials. The left panel reports results obtained for models—random forest, SVM, RuleFit, and SR4-fit , whereas the right panel shows results with models LASSO regression, decision tree, and XGBoost.
Figure 57 : Line plot comparison of model performance metrics for housing data across 30 trials. The left panel reports results obtained for models—random forest, SVM, RuleFit, and SR4-fit , whereas the right panel shows results with models LASSO regression, decision tree, and XGBoost.
Figure 58 : Line plot comparison of model performance metrics for machine data across 30 trials. The left panel reports results obtained for models—random forest, SVM, RuleFit, and SR4-fit , whereas the right panel shows results with models LASSO regression, decision tree, and XGBoost.
Figure 59 : Line plot comparison of model performance metrics for MPG data across 30 trials. The left panel reports results obtained for models—random forest, SVM, RuleFit, and SR4-fit , whereas the right panel shows results with models LASSO regression, decision tree, and XGBoost.
Figure 60 : Line plot comparison of model performance metrics for ozone data across 30 trials. The left panel reports results obtained for models—random forest, SVM, RuleFit, and SR4-fit , whereas the right panel shows results with models LASSO regression, decision tree, and XGBoost.
Figure 61 : Line plot comparison of model performance metrics for prostate data across 30 trials. The left panel reports results obtained for models—random forest, SVM, RuleFit, and SR4-fit , whereas the right panel shows results with models LASSO regression, decision tree, and XGBoost.
Figure 62 : Violin plot comparison of models (random forest, SVM, RuleFit, and SR4-fit) performance across multiple public regression datasets for different prediction metrics.
Figure 63 : Violin plot comparison of models (LASSO regression, decision tree, and XGBoost) performance across multiple public regression datasets for different prediction metrics.
Dataset
Random Forest
RuleFit
SR4-Fit
Decision Tree
Abalone
0.0001 ± 0.0000
0.5000 ± 0.0000
0.5011 ± 0.0028
0.2506 ± 0.1684
Bone
0.0012 ± 0.0003
0.2727 ± 0.0000
0.3710 ± 0.0326
0.0230 ± 0.0228
Diabetes
0.0001 ± 0.0001
0.5556 ± 0.0000
0.5848 ± 0.0156
0.1425 ± 0.0970
Housing
0.0005 ± 0.0002
0.6190 ± 0.0000
0.5880 ± 0.0143
0.0428 ± 0.0668
Machine
0.0090 ± 0.0020
0.4138 ± 0.0141
0.6146 ± 0.0167
0.0169 ± 0.0205
MPG
0.0005 ± 0.0001
0.5000 ± 0.0000
0.3743 ± 0.0095
1.0000 ± 0.0000
Appendix
Table 21 : Average Dice–Sørensen Index ± standard deviation across 30 trials for public regression datasets. Bold indicates the best-performing method per dataset.
Dataset
Random Forest
RuleFit
SR4-Fit
Decision Tree
Abalone
8666.53 ± 101.33
16.00 ± 0.00
15.96 ± 0.18
1.00 ± 0.00
Bone
8151.33 ± 298.36
11.00 ± 0.00
2.73 ± 0.44
1.03 ± 0.18
Diabetes
7823.36 ± 209.19
18.00 ± 0.00
17.00 ± 0.69
1.00 ± 0.00
Housing
7244.06 ± 142.93
21.00 ± 0.00
21.90 ± 0.95
3.06 ± 1.08
Machine
4166.13 ± 207.00
21.80 ± 1.5177
14.66 ± 0.84
2.66 ± 0.80
MPG
9090.80 ± 205.65
16.00 ± 0.00
21.40 ± 1.13
1.00 ± 0.00
Appendix
Table 22 : Average number of rules ± standard deviation across 30 trials for public regression datasets. Lower values indicate more compact models.
Dataset
Random Forest
RuleFit
SR4-Fit
Decision Tree
Abalone
6.54 ± 0.02
2.00 ± 0.00
1.99 ± 0.01
1.00 ± 0.00
Bone
5.95 ± 0.03
2.45 ± 0.00
2.24 ± 0.15
1.73 ± 0.98
Diabetes
5.88 ± 0.02
1.88 ± 0.00
1.82 ± 0.05
1.00 ± 0.00
Housing
5.89 ± 0.02
1.76 ± 0.00
2.22 ± 0.07
3.00 ± 0.00
Machine
5.62 ± 0.05
2.61 ± 0.14
1.75 ± 0.07
2.72 ± 0.58
MPG
5.98 ± 0.02
2.00 ± 0.00
2.87 ± 0.06
1.00 ± 0.00
Appendix
Table 23 : Average rule complexity ± standard deviation across 30 trials for public regression datasets. Bold indicates less complex rules
Interpretable machine learning is essential in high-stakes domains where decision-making requires accountability, transparency, and trust. While rule-based models offer global and exact interpretability, learning rule sets that simultaneously achieve high predictive performance and low, human-understandable complexity remains challenging. To address this, we introduce TT-Sparse, a flexible neural building block that leverages differentiable truth tables as nodes to learn sparse, effective connections. A key contribution of our approach is a new soft TopK operator with straight-through estimation for learning discrete, cardinality-constrained feature selection in an end-to-end differentiable manner. Crucially, the forward pass remains sparse, enabling efficient computation and exact symbolic rule extraction. As a result, each node (and the entire model) can be transformed exactly into compact, globally interpretable DNF/CNF Boolean formulas via Quine-McCluskey minimization. Extensive empirical results across 28 datasets spanning binary, multiclass, and regression tasks show that the learned sparse rules exhibit superior predictive performance with lower complexity compared to existing state-of-the-art methods.
Hans Farrell Soegeng, Sarthak Ketanbhai Modi, Thomas Peyrin
School of Physical and Mathematical Sciences, Nanyang Technological University, Singapore.
Machine learning models achieve high predictive accuracy in regression tasks, but their deployment in safety-critical and regulated domains requires interpretability. While fuzzy rule-based systems offer transparent, linguistically explicit interpretable models, Mamdani-style fuzzy regression remains underrepresented in modern machine learning software libraries. This paper presents an interpretable regression extension for the Ex-Fuzzy library, enabling Mamdani fuzzy inference with scalar consequents learned directly from data. For this, a target-aware partition initialisation strategy based on Fuzzy C-Means clustering is introduced, in which linguistic variables are derived from an augmented input-output space to emphasise output-relevant regions of the feature space. The proposed extension is evaluated on ten regression datasets from the KEEL repository, comparing Gaussian and trapezoidal partition strategies against standard baselines including linear regression, multilayer perceptron, and random forests. Experimental results show that Gaussian partitions consistently outperform uniform trapezoidal partitions, achieving a mean coefficient of determination of approximately 0.86 while producing compact rule bases of 10-15 human-readable rules. The proposed implementation provides a transparent and competitive alternative to black-box regression models, supporting practical interpretability with competitive predictive performance.
Cayan Deniz Kucuktopana, Javier Fumanal-Idocin, Richard Pitts +1
School of Computer Science and Electronic Engineering University of Essex Colchester, United Kingdom
Recent work has shown that well-optimized individual decision trees can match complex black box models in some settings, primarily in noisy domains. For the remaining settings, however, complex ensembled compositions of trees often achieve higher accuracy at the cost of interpretability, leaving practitioners with difficult modeling decisions along an accuracy-interpretability tradeoff. Ideally, we would like to classify as much of the data as possible with one or a small number of trees, achieving interpretability for most samples while maintaining state-of-the-art accuracy. We introduce Multistage Defer Trees: a sequence of sparse decision trees that each make predictions for most samples, while deferring a small proportion to the next tree in the sequence or, ultimately, to a black box. We demonstrate that we can train this model class to match the performance of complex tree-based ensembles while routing most samples through only one or a small number of sparse decision trees. We discuss a range of techniques for training these models while maintaining simplicity. Our method expands the accuracy--interpretability frontier in settings where single-tree methods remain insufficient, demonstrating that even when complex models are necessary, they need not be fully opaque.
Zakk Heile, Hayden McTavish, Margo Seltzer +1
Department of Computer Science · Duke University · Durham, USA +2