In many high-stakes applications, machine learning is dominated by black-box models that require post hoc explanations to justify their predictions. These explanations are often unreliable because they do not reflect the model's actual computations, limiting accountability and trust. A natural alternative is to use models that are interpretable by design. However, existing rule-based approaches, such as RuleFit and decision trees, while transparent, often lack stability and predictive strength, reinforcing a perceived trade-off between traditional performance measures and model understandability. To address this, we propose Sparse Relaxed Regularized Regression Rule-Fit (SR4-Fit), an intrinsically interpretable algorithm for both classification and regression that produces compact and stable rule sets without sacrificing performance. Using demographic data from the U.S. Census Bureau's American Community Survey, SR4-Fit predicts U.S. House election outcomes with high accuracy and interpretability while uncovering demographic interactions missed by black-box models. We further validate SR4-Fit across fourteen benchmark datasets (six classification and eight regression), where it outperforms existing rule-based methods, including RuleFit and decision trees in terms of accuracy, stability, and compactness while remaining competitive with black-box models in predictivity. These results demonstrate that interpretability and predictive reliability need not be mutually exclusive, offering a practical and transparent alternative for high-stakes decision-making.
Figures & tables
Figure 1 : Conceptual illustration of the predictivity–understandability trade-off across common machine learning models. SR4-Fit is positioned closer to the ideal region of high predictivity and high understandability, reflecting its aim to balance both criteria more effectively than existing rule-based approaches.
Figure 2 : Overview of the SR4-Fit pipeline (including classification and regression).
β=argminβ21∥y−Zβ∥22+2κ∥β−w∥22
Algorithm 2 SR4-Fit Regression Algorithm
Figure 3 : Aggregate Pareto plots for the U.S. House election classification and regression tasks, showing mean predictivity (F1 score for classification, left; R2 score for regression, right) against Rule Understandability Score (RUS). The classification panel averages results across all four demographic datasets and both feature configurations (with and without party percentages). The regression panel averages results across all four datasets and all four experimental configurations (DEM as label with and without REP percentage; REP as label with and without DEM percentage).
Figure 4 : Aggregate Pareto plots for the public benchmark classification and regression tasks, showing mean predictivity (F1 score for classification, left; R2 score for regression, right) against Rule Understandability Score (RUS). The classification panel averages results across the six public classification benchmark datasets. The regression panel averages results across the eight public regression benchmark datasets.
Table 1 : Average enhanced interpretability score (E-IPS) ± standard deviation across 30 trials for demographic classification datasets. Bold indicates the best-performing method per dataset.
Figure 5 : Distribution of enhanced interpretability scores (E-IPS) across demographic classification datasets. The left panel reports results obtained using feature sets that include party percentage information, whereas the right panel shows results with party percentages removed.
Table 2 : Average enhanced interpretability score (E-IPS) ± standard deviation across 30 trials for demographic regression datasets. Results are reported for predicting Democratic vote percentage, with and without Republican percentage included as a feature. Bold indicates the best-performing method per dataset.
Figure 6 : Distribution of enhanced interpretability scores (E-IPS) across demographic regression datasets when predicting Democratic vote percentage. Results are shown for models trained with all party-related features (left) and for models trained with Republican percentage features removed (right).
Table 3 : Average enhanced interpretability score (E-IPS) ± standard deviation across 30 trials for demographic regression datasets. Results are reported for predicting Republican vote percentage, with and without Democratic percentage included as a feature. Bold indicates the best-performing method per dataset.
Figure 7 : Distribution of enhanced interpretability scores (E-IPS) across demographic regression datasets when predicting Republican vote percentage. Results are shown for models trained with all party-related features (left) and for models trained with Democratic percentage features removed (right).
Dataset
Random Forest
RuleFit
SR4-Fit
Decision Tree
Breast Cancer
0.4350 ± 0.1433
0.4793 ± 0.1851
0.4931 ± 0.1875
0.4477 ± 0.2044
E. coli
0.4505 ± 0.1475
0.6085 ± 0.1087
0.6141 ± 0.1093
0.3624 ± 0.1494
Page Blocks
0.5386 ± 0.1296
0.4813 ± 0.1322
0.5308 ± 0.1381
0.5405 ± 0.1391
Pima Indians
0.4644 ± 0.1330
0.3596 ± 0.3061
0.3208 ± 0.3011
0.3862 ± 0.1781
Vehicle
0.4958 ± 0.1642
0.4175 ± 0.1314
0.4099 ± 0.1262
0.5338 ± 0.1570
Yeast
0.4927 ± 0.1426
0.6313 ± 0.1234
0.6512 ± 0.1281
0.4017 ± 0.1596
Table 4 : Average enhanced interpretability score (E-IPS) ± standard deviation across 30 trials for public classification datasets. Bold indicates the highest average E-IPS per dataset.
Figure 8 : Violin plot comparison of enhanced interpretability scores for all rule-based models across standard public classification datasets.
Dataset
Random Forest
RuleFit
SR4-Fit
Decision Tree
Abalone
0.4890 ± 0.1166
0.4824 ± 0.0649
0.1444 ± 0.1525
0.5524 ± 0.1113
Bone
0.4723 ± 0.1565
0.5756 ± 0.0549
0.3263 ± 0.3564
0.6124 ± 0.2261
Diabetes
0.4019 ± 0.1627
0.6474 ± 0.0588
0.6176 ± 0.2149
0.5338 ± 0.0878
Housing
0.4965 ± 0.1320
0.5330 ± 0.0613
0.5425 ± 0.1730
0.4301 ± 0.1264
Machine
0.5615 ± 0.1073
0.5410 ± 0.1868
0.5588 ± 0.1688
0.4165 ± 0.1849
MPG
0.4738 ± 0.1402
0.5391 ± 0.0615
0.5393 ± 0.2905
0.5655 ± 0.0507
Table 5 : Average enhanced interpretability score (E-IPS) ± standard deviation across 30 trials for public regression datasets. Bold indicates the highest average E-IPS per dataset.
Figure 9 : Violin plot comparison of enhanced interpretability scores for all rule-based models across standard public regression datasets.
Appendix figures & tables72 assets
Supplementary material from the paper’s appendix.
Appendix
Table 6 : Wilcoxon test W-statistic and p-value comparison between RuleFit and SR4-Fit over 30 trials across demographic classification datasets. Statistically significant differences ( p<0.05 ) are shown in bold .
Table 7 : Wilcoxon signed-rank test W-statistic and p-value comparison between RuleFit and SR4-Fit over 30 trials across demographic datasets for Democratic percentage label regression. Statistically significant differences ( p<0.05 ) are shown in bold .
Table 8 : Wilcoxon signed-rank test W-statistic and p-value comparison between RuleFit and SR4-Fit over 30 trials across demographic datasets for Republican percentage label regression. Statistically significant differences ( p<0.05 ) are shown in bold .
Figure 10 : Line plot comparison of model performance metrics for minimum data across 30 trials. The left panel reports results obtained using feature sets that include prior party voting percentage information, whereas the right panel shows results with prior party voting percentages removed.
Figure 11 : Line plot comparison of model (logistic regression, Decision Tree, XG boost) performance metrics for minimum data across 30 trials. The left panel reports results obtained using feature sets that include prior party voting percentage information, whereas the right panel shows results with prior party voting percentages removed.
Figure 12 : Line plot comparison of model performance metrics for standard data across 30 trials. The left panel reports results obtained using feature sets that include prior party voting percentage information, whereas the right panel shows results with prior party voting percentages removed.
Figure 13 : Line plot comparison of model (logistic regression, Decision Tree, XG boost) performance metrics for standard data across 30 trials. The left panel reports results obtained using feature sets that include prior party voting percentage information, whereas the right panel shows results with prior party voting percentages removed.
Figure 14 : Line plot comparison of model performance metrics for expanded data across 30 trials. The left panel reports results obtained using feature sets that include prior party voting percentage information, whereas the right panel shows results with prior party voting percentages removed.
Figure 15 : Line plot comparison of model (logistic regression, Decision Tree, XG boost) performance metrics for expanded data across 30 trials. The left panel reports results obtained using feature sets that include prior party voting percentage information, whereas the right panel shows results with prior party voting percentages removed.
Figure 16 : Line plot comparison of model performance metrics for previous data across 30 trials. The left panel reports results obtained using feature sets that include prior party voting percentage information, whereas the right panel shows results with prior party voting percentages removed.
Figure 17 : Line plot comparison of model (logistic regression, Decision Tree, XG boost) performance metrics for previous data across 30 trials. The left panel reports results obtained using feature sets that include prior party voting percentage information, whereas the right panel shows results with prior party voting percentages removed.
Figure 18 : Violin plot comparison of model performance metrics across multiple datasets for the election classification task that includes prior party voting percentage information.
Figure 19 : Violin plot comparison of model (logistic regression, Decision Tree, XG boost) performance metrics across multiple datasets for the election classification task that includes prior party voting percentage information.
Figure 20 : Violin plot comparison of model performance metrics across multiple datasets for the election classification task that does not include prior party voting percentage information.
Figure 21 : Violin plot comparison of model (logistic regression, Decision Tree, XG boost) performance metrics across multiple datasets for the election classification task that does not include prior party voting percentage information.
Table 9 : Average Dice–Sørensen Index ± standard deviation across 30 trials for election classification datasets. Bold indicates the best-performing method per dataset.
Table 10 : Average number of rules ± standard deviation across 30 trials for election classification datasets. Bold indicates the most compact model.
Table 11 : Average rule complexity (number of conditions per rule) ± standard deviation for election classification datasets across 30 trials. Bold indicates less complex rules.
Figure 22 : Line plot comparison of model performance metrics for minimum data across 30 trials with DEM percentage as label. The left panel reports results obtained with party percentage (REP) included, whereas the right panel shows results with party percentage (REP) not included.
Figure 23 : Line plot comparison of model (LASSO regression, Decision Tree, XG boost) performance metrics for minimum data across 30 trials. The left panel reports results obtained with party percentage (REP) included, whereas the right panel shows results with party percentage (REP) not included.
Figure 24 : Line plot comparison of model performance metrics for standard data across 30 trials with DEM percentage as label. The left panel reports results obtained with party percentage (REP) included, whereas the right panel shows results with party percentage (REP) not included.
Figure 25 : Line plot comparison of model (LASSO regression, Decision Tree, XG boost) performance metrics for standard data across 30 trials. The left panel reports results obtained with party percentage (REP) included, whereas the right panel shows results with party percentage (REP) not included.
Figure 26 : Line plot comparison of model performance metrics for expanded data across 30 trials with DEM percentage as label. The left panel reports results obtained with party percentage (REP) included, whereas the right panel shows results with party percentage (REP) not included.
Figure 27 : Line plot comparison of model (LASSO regression, Decision Tree, XG boost) performance metrics for expanded data across 30 trials. The left panel reports results obtained with party percentage (REP) included, whereas the right panel shows results with party percentage (REP) not included.
Figure 28 : Line plot comparison of model performance metrics for previous data across 30 trials with DEM percentage as label. The left panel reports results obtained with party percentage (REP) included, whereas the right panel shows results with party percentage (REP) not included.
Figure 29 : Line plot comparison of model (LASSO regression, Decision Tree, XG boost) performance metrics for previous data across 30 trials. The left panel reports results obtained with party percentage (REP) included, whereas the right panel shows results with party percentage (REP) not included.
Figure 30 : Violin plot comparison of model performance metrics across multiple datasets for the election regression task that has DEM percentage as a label and includes party percentage (REP).
Figure 31 : Violin plot comparison of model (LASSO regression, Decision Tree, XG boost) performance metrics across multiple datasets for the election regression task that has DEM percentage as a label and includes party percentage (REP).
Figure 32 : Violin plot comparison of model performance metrics across multiple datasets for the election regression task that has DEM percentage as a label and does not include party percentage (REP).
Figure 33 : Violin plot comparison of model (LASSO regression, Decision Tree, XG boost) performance metrics across multiple datasets for the election regression task that has DEM percentage as a label and does not include party percentage (REP).
Table 12 : Average Dice–Sørensen Index ± standard deviation across 30 trials for the election regression dataset where DEM percentage is the label.
Table 13 : Average number of rules ± standard deviation across 30 trials for the election regression dataset, where DEM percentage is the label.
Table 14 : Average rule complexity ± standard deviation across 30 trials for the election regression dataset where DEM percentage is the label.
Figure 34 : Line plot comparison of model performance metrics for minimum data across 30 trials with REP percentage as label. The left panel reports results obtained with party percentage (DEM) included, whereas the right panel shows results with party percentage (DEM) not included.
Figure 35 : Line plot comparison of model (LASSO regression, decision tree, XG boost) performance metrics for minimum data across 30 trials. The left panel reports results obtained with party percentage (DEM) included, whereas the right panel shows results with party percentage (DEM) not included.
Figure 36 : Line plot comparison of model performance metrics for standard data across 30 trials with REP percentage as label. The left panel reports results obtained with party percentage (DEM) included, whereas the right panel shows results with party percentage (DEM) not included.
Figure 37 : Line plot comparison of model (LASSO regression, decision tree, XG boost) performance metrics for standard data across 30 trials. The left panel reports results obtained with party percentage (DEM) included, whereas the right panel shows results with party percentage (DEM) not included.
Figure 38 : Line plot comparison of model performance metrics for expanded data across 30 trials with REP percentage as label. The left panel reports results obtained with party percentage (DEM) included, whereas the right panel shows results with party percentage (DEM) not included.
Figure 39 : Line plot comparison of model (LASSO regression, decision tree, XG boost) performance metrics for expanded data across 30 trials. The left panel reports results obtained with party percentage (DEM) included, whereas the right panel shows results with party percentage (DEM) not included.
Figure 40 : Line plot comparison of model performance metrics for previous data across 30 trials with REP percentage as label. The left panel reports results obtained with party percentage (DEM) included, whereas the right panel shows results with party percentage (DEM) not included.
Figure 41 : Line plot comparison of model (LASSO regression, decision tree, XG boost) performance metrics for previous data across 30 trials. The left panel reports results obtained with party percentage (DEM) included, whereas the right panel shows results with party percentage (DEM) not included.
Figure 42 : Violin plot comparison of model performance metrics across multiple datasets for the election regression task that has REP percentage as a label and includes party percentage (DEM).
Figure 43 : Violin plot comparison of model (LASSO regression, decision tree, XG boost) performance metrics across multiple datasets for the election regression task that has REP percentage as a label and includes party percentage (DEM).
Figure 44 : Violin plot comparison of model performance metrics across multiple datasets for the election regression task that has REP percentage as a label and does not include party percentage (DEM).
Figure 45 : Violin plot comparison of model (LASSO regression, decision tree, XG boost) performance metrics across multiple datasets for the election regression task that has REP percentage as a label and does not include party percentage (DEM).
Table 15 : Average Dice–Sørensen Index ± standard deviation across 30 trials for the election regression dataset where REP percentage is the label.
Table 16 : Average number of rules ± standard deviation across 30 trials for the election regression dataset, where REP percentage is the label.
Table 17 : Average rule complexity ± standard deviation across 30 trials for the election regression dataset, where REP percentage is the label.
Figure 46 : Line plot comparison of model performance metrics for breast cancer data across 30 trials. The left panel reports results obtained for models—random forest, SVM, rulefit, and SR4-fit , whereas the right panel shows results with models logistic regression, decision tree, and XGboost.
Figure 47 : Line plot comparison of model performance metrics for E. coli data across 30 trials. The left panel reports results obtained for models—random forest, SVM, rulefit, and SR4-fit , whereas the right panel shows results with models logistic regression, decision tree, and XGboost.
Figure 48 : Line plot comparison of model performance metrics for page blocks data across 30 trials. The left panel reports results obtained for models—random forest, SVM, rulefit, and SR4-fit , whereas the right panel shows results with models logistic regression, decision tree, and XGboost.
Figure 49 : Line plot comparison of model performance metrics for Pima Indians data across 30 trials. The left panel reports results obtained for models—random forest, SVM, rulefit, and SR4-fit , whereas the right panel shows results with models logistic regression, decision tree, and XGboost.
Figure 50 : Line plot comparison of model performance metrics for vehicle data across 30 trials. The left panel reports results obtained for models—random forest, SVM, rulefit, and SR4-fit , whereas the right panel shows results with models logistic regression, decision tree, and XGboost.
Figure 51 : Line plot comparison of model performance metrics for yeast data across 30 trials. The left panel reports results obtained for models—random forest, SVM, rulefit, and SR4-fit , whereas the right panel shows results with models logistic regression, decision tree, and XGboost.
Figure 52 : Violin plot comparison of model (random forest, SVM, RuleFit, and SR4-fit) performance across multiple standard benchmark classification datasets for different prediction metrics.
Figure 53 : Violin plot comparison of model (logistic regression, decision tree, and XGBoost) performance across multiple standard benchmark classification datasets for different prediction metrics.
Dataset
Random Forest
RuleFit
SR4-Fit
Decision Tree
Breast Cancer
0.3307 ± 0.0186
0.4449 ± 0.0199
0.4449 ± 0.0199
0.0558 ± 0.0540
E. coli
0.1160 ± 0.0150
0.1762 ± 0.0140
0.1762 ± 0.0140
0.0414 ± 0.0347
Page Blocks
0.1090 ± 0.0155
0.3232 ± 0.0294
0.3232 ± 0.0294
0.0135 ± 0.0090
Pima Indians
0.0454 ± 0.0078
0.5153 ± 0.0115
0.5153 ± 0.0115
0.0021 ± 0.0032
Vehicle
0.1993 ± 0.0146
0.5231 ± 0.0272
0.5231 ± 0.0272
0.0593 ± 0.0360
Yeast
0.1182 ± 0.0146
0.1807 ± 0.0182
0.1807 ± 0.0182
0.0121 ± 0.0098
Appendix
Table 18 : Average Dice–Sørensen Index ± standard deviation across 30 trials for public classification datasets. Bold indicates the best-performing method per dataset.
Dataset
Random Forest
RuleFit
SR4-Fit
Decision Tree
Breast Cancer
104.43 ± 11.58
21.13 ± 1.11
21.13 ± 1.11
9.10 ± 0.66
E. coli
117.17 ± 13.48
57.00 ± 0.00
57.00 ± 0.00
11.03 ± 0.67
Page Blocks
50.27 ± 8.77
43.10 ± 2.15
43.10 ± 2.15
18.23 ± 0.89
Pima Indians
37.47 ± 8.19
15.67 ± 0.48
15.67 ± 0.48
20.77 ± 1.96
Vehicle
147.37 ± 11.09
40.07 ± 0.78
40.07 ± 0.78
18.57 ± 1.33
Yeast
57.00 ± 8.87
57.00 ± 0.00
57.00 ± 0.00
24.07 ± 1.41
Appendix
Table 19 : Average number of rules ± standard deviation across 30 trials for public classification datasets. Lower values indicate more compact models.
Dataset
Random Forest
RuleFit
SR4-Fit
Decision Tree
Breast Cancer
1.99 ± 0.01
2.97 ± 0.12
2.97 ± 0.12
3.42 ± 0.15
E. coli
1.97 ± 0.01
3.55 ± 0.01
3.55 ± 0.01
3.60 ± 0.08
Page Blocks
1.97 ± 0.01
2.47 ± 0.04
2.47 ± 0.04
4.41 ± 0.05
Pima Indians
1.96 ± 0.01
2.28 ± 0.07
2.28 ± 0.07
4.61 ± 0.08
Vehicle
1.99 ± 0.01
2.05 ± 0.04
2.05 ± 0.04
4.52 ± 0.09
Yeast
1.96 ± 0.02
2.66 ± 0.03
2.66 ± 0.03
4.69 ± 0.06
Appendix
Table 20 : Average rule complexity (number of conditions per rule) ± standard deviation across 30 trials for public classification datasets. Bold indicates less complex rules.
Figure 54 : Line plot comparison of model performance metrics for abalone data across 30 trials. The left panel reports results obtained for models—random forest, SVM, RuleFit, and SR4-fit , whereas the right panel shows results with models LASSO regression, decision tree, and XGBoost.
Figure 55 : Line plot comparison of model performance metrics for bone data across 30 trials. The left panel reports results obtained for models—random forest, SVM, RuleFit, and SR4-fit, whereas the right panel shows results with models LASSO regression, decision tree, and XGBoost.
Figure 56 : Line plot comparison of model performance metrics for diabetes data across 30 trials. The left panel reports results obtained for models—random forest, SVM, RuleFit, and SR4-fit , whereas the right panel shows results with models LASSO regression, decision tree, and XGBoost.
Figure 57 : Line plot comparison of model performance metrics for housing data across 30 trials. The left panel reports results obtained for models—random forest, SVM, RuleFit, and SR4-fit , whereas the right panel shows results with models LASSO regression, decision tree, and XGBoost.
Figure 58 : Line plot comparison of model performance metrics for machine data across 30 trials. The left panel reports results obtained for models—random forest, SVM, RuleFit, and SR4-fit , whereas the right panel shows results with models LASSO regression, decision tree, and XGBoost.
Figure 59 : Line plot comparison of model performance metrics for MPG data across 30 trials. The left panel reports results obtained for models—random forest, SVM, RuleFit, and SR4-fit , whereas the right panel shows results with models LASSO regression, decision tree, and XGBoost.
Figure 60 : Line plot comparison of model performance metrics for ozone data across 30 trials. The left panel reports results obtained for models—random forest, SVM, RuleFit, and SR4-fit , whereas the right panel shows results with models LASSO regression, decision tree, and XGBoost.
Figure 61 : Line plot comparison of model performance metrics for prostate data across 30 trials. The left panel reports results obtained for models—random forest, SVM, RuleFit, and SR4-fit , whereas the right panel shows results with models LASSO regression, decision tree, and XGBoost.
Figure 62 : Violin plot comparison of models (random forest, SVM, RuleFit, and SR4-fit) performance across multiple public regression datasets for different prediction metrics.
Figure 63 : Violin plot comparison of models (LASSO regression, decision tree, and XGBoost) performance across multiple public regression datasets for different prediction metrics.
Dataset
Random Forest
RuleFit
SR4-Fit
Decision Tree
Abalone
0.0001 ± 0.0000
0.5000 ± 0.0000
0.5011 ± 0.0028
0.2506 ± 0.1684
Bone
0.0012 ± 0.0003
0.2727 ± 0.0000
0.3710 ± 0.0326
0.0230 ± 0.0228
Diabetes
0.0001 ± 0.0001
0.5556 ± 0.0000
0.5848 ± 0.0156
0.1425 ± 0.0970
Housing
0.0005 ± 0.0002
0.6190 ± 0.0000
0.5880 ± 0.0143
0.0428 ± 0.0668
Machine
0.0090 ± 0.0020
0.4138 ± 0.0141
0.6146 ± 0.0167
0.0169 ± 0.0205
MPG
0.0005 ± 0.0001
0.5000 ± 0.0000
0.3743 ± 0.0095
1.0000 ± 0.0000
Appendix
Table 21 : Average Dice–Sørensen Index ± standard deviation across 30 trials for public regression datasets. Bold indicates the best-performing method per dataset.
Dataset
Random Forest
RuleFit
SR4-Fit
Decision Tree
Abalone
8666.53 ± 101.33
16.00 ± 0.00
15.96 ± 0.18
1.00 ± 0.00
Bone
8151.33 ± 298.36
11.00 ± 0.00
2.73 ± 0.44
1.03 ± 0.18
Diabetes
7823.36 ± 209.19
18.00 ± 0.00
17.00 ± 0.69
1.00 ± 0.00
Housing
7244.06 ± 142.93
21.00 ± 0.00
21.90 ± 0.95
3.06 ± 1.08
Machine
4166.13 ± 207.00
21.80 ± 1.5177
14.66 ± 0.84
2.66 ± 0.80
MPG
9090.80 ± 205.65
16.00 ± 0.00
21.40 ± 1.13
1.00 ± 0.00
Appendix
Table 22 : Average number of rules ± standard deviation across 30 trials for public regression datasets. Lower values indicate more compact models.
Dataset
Random Forest
RuleFit
SR4-Fit
Decision Tree
Abalone
6.54 ± 0.02
2.00 ± 0.00
1.99 ± 0.01
1.00 ± 0.00
Bone
5.95 ± 0.03
2.45 ± 0.00
2.24 ± 0.15
1.73 ± 0.98
Diabetes
5.88 ± 0.02
1.88 ± 0.00
1.82 ± 0.05
1.00 ± 0.00
Housing
5.89 ± 0.02
1.76 ± 0.00
2.22 ± 0.07
3.00 ± 0.00
Machine
5.62 ± 0.05
2.61 ± 0.14
1.75 ± 0.07
2.72 ± 0.58
MPG
5.98 ± 0.02
2.00 ± 0.00
2.87 ± 0.06
1.00 ± 0.00
Appendix
Table 23 : Average rule complexity ± standard deviation across 30 trials for public regression datasets. Bold indicates less complex rules