Neural networks (NNs) achieve outstanding performance in many domains; however, their decision processes are often opaque and their inference can be computationally expensive in resource-constrained environments. We recently proposed Differentiable Logic Networks (DLNs) to address these issues for tabular classification based on relaxing discrete logic into a differentiable form, thereby enabling gradient-based learning of networks built from binary logic operations. DLNs offer interpretable reasoning and substantially lower inference cost. We extend the DLN framework to supervised tabular regression. We first redesign the final output layer (the SumLayer) to support continuous targets. More critically, we find the original two-phase training procedure used for classification is suboptimal for regression, and thus develop a unified, single-stage optimization procedure. We also demonstrate that temperature annealing of the network's differentiable relaxations is decisive for achieving stable convergence and high accuracy. We evaluate the resulting model on 15 public regression benchmarks, comparing it with modern neural networks and classical regression baselines. Regression DLNs match or exceed baseline accuracy while preserving interpretability and fast inference. Our results show that DLNs are a viable, cost-effective alternative for regression tasks, especially where model transparency and computational efficiency are important.
Figures & tables
Figure 1: A simplified regression DLN example. Continuous input features are first binarized by a ThresholdLayer, producing a binary vector. This vector is then processed by successive LogicLayers composed of two-input Boolean operators. The activations of the final LogicLayer are combined by a SumLayer that computes a weighted sum, yielding the real-valued prediction.
Figure 2: Training workflow for a regression DLN.
ThresholdLayer
LogicLayer
SumLayer
Trainable parameters
bias b∈Rin_dim slope s∈Rin_dim
logic fn weight W∈Rout_dim×16 link a weight U∈Rout_dim×in_dim link b weight V∈Rout_dim×in_dim
Table 1: Summary of DLN trainable parameters and feedforward functions during unified training and during inference. We use x to denote input and y to denote the output of each layer.
Figure 3: Illustration of the training process. The network learns neuron functions and connections simultaneously: Key details are highlighted for clarity.
Dataset
# Train
# Test
# Cont.
# Cate.
Source
ID
Abalone
3132
1045
7
2
UCI
1
Airfoil
1127
376
5
0
UCI
291
Bike
13032
4345
6
14
UCI
275
CCPP
7145
2382
4
0
UCI
294
Concrete
753
252
8
0
UCI
165
Electrical
7500
2500
12
0
UCI
471
Table 2: Characteristics of the datasets after preprocessing
Figure 4: Comparison of R2 and the number of operations required for inference across models. We plot results for 32, 64, and 128 hyperparameter-search trials and draw the Pareto frontier for each case. DLN is Pareto-optimal in all three. XGBoost is the best-performing model in terms of accuracy. The vanilla DLN ensembles extend the Pareto frontier, outperforming random forest, a model that uses a similar ensembling framework, on both accuracy and inference cost.
Linear
Ridge
Lasso
KNN
DT
AB
RF
XGB
SVR
MLP
DLN
Abalone
0.534 ± 0.024
0.534 ± 0.024
0.534 ± 0.024
0.517 ± 0.013
0.475 ± 0.018
0.463 ± 0.027
0.547 ± 0.019
0.554 ± 0.018
0.551 ± 0.018
0.570 ± 0.024
0.526 ± 0.026
Airfoil
0.516 ± 0.034
0.516 ± 0.033
0.516 ± 0.033
0.874 ± 0.026
0.793 ± 0.052
0.743 ± 0.020
0.868 ± 0.016
0.952 ± 7.4e-03
0.822 ± 0.024
0.943 ± 5.7e-03
0.889 ± 0.013
Bike
0.403 ± 0.012
0.403 ± 0.012
0.403 ± 0.012
0.659 ± 0.012
0.904 ± 0.011
0.660 ± 0.017
0.870 ± 7.1e-03
0.956 ± 3.2e-03
0.645 ± 0.014
0.939 ± 3.0e-03
0.916 ± 6.4e-03
CCPP
0.929 ± 2.4e-03
0.929 ± 2.4e-03
0.929 ± 2.4e-03
0.957 ± 2.0e-03
0.942 ± 2.8e-03
0.920 ± 3.8e-03
0.947 ± 4.6e-03
0.966 ± 1.3e-03
0.946 ± 2.2e-03
0.948 ± 2.4e-03
0.943 ± 2.3e-03
Concrete
0.593 ± 0.034
0.593 ± 0.034
0.593 ± 0.034
0.709 ± 0.038
0.795 ± 0.029
0.791 ± 0.015
0.877 ± 0.018
0.931 ± 0.010
0.869 ± 0.014
0.882 ± 0.022
0.888 ± 0.020
Electrical
0.647 ± 0.014
0.647 ± 0.014
0.647 ± 0.013
0.798 ± 4.0e-03
0.755 ± 0.012
0.826 ± 8.4e-03
0.843 ± 0.013
0.955 ± 2.7e-03
0.962 ± 1.7e-03
0.968 ± 2.0e-03
0.936 ± 3.6e-03
Table 3: Average test R2 across 10 random seeds and 128 hyperparameter trials per model
Figure 5: Distribution of R2 across datasets for each model, averaged over 10 random seeds and 128 hyperparameter-search trials.
Figure 6: Pearson correlation matrix of model R2 scores, averaged over 10 random seeds and 128 hyperparameter-search trials.
32 trials
64 trials
128 trials
Linear
0.614
0.614 →
0.614 →
Ridge
0.615
0.615 →
0.615 →
Lasso
0.615
0.615 →
0.615 →
KNN
0.726
0.727 ↑
0.727 →
DT
0.755
0.771 ↑
0.777 ↑
AB
0.722
0.723 ↑
0.725 ↑
Table 4: Mean R2 scores across HPO budgets
Linear
Ridge
Lasso
KNN
DT
AB
RF
XGB
SVR
MLP
DLN
Abalone
11.2K
11.2K
10.7K
38.5M
1.19K
206K
240K
432K
383M
829K
41.2K
Airfoil
6.22K
6.22K
6.22K
9.66M
2.12K
295K
174K
853K
109M
1.71M
32.6K
Bike
24.9K
24.9K
21.9K
271M
2.85K
158K
177K
1.01M
1.30G
12.9M
50.0K
CCPP
4.98K
4.98K
4.98K
38.9M
2.01K
199K
201K
960K
504M
1.22M
22.5K
Concrete
9.95K
9.95K
8.71K
9.46M
2.05K
293K
161K
678K
67.6M
2.57M
60.6K
Electrical
14.9K
14.9K
13.3K
89.5M
2.36K
296K
191K
840K
1.03G
4.65M
91.1K
Table 5: Average number of basic gate-level logic operations required for inference across 10 random seeds and 128 hyperparameter trials, assuming FP16 for floating-point and INT16 for integer arithmetic
Batch Size
XGB
MLP
DLN
Batch = 1
227 μs
39.8 μs
12.6 μs
Batch = 32
13.8 μs
1.62 μs
1.11 μs
Table 6: Per-sample inference latency (in microseconds) on a single-threaded CPU. Values are the geometric mean (across 15 datasets) of the per-dataset average latencies (mean over 10 seeds). Lower is better.
Figure 7: Decision process learned by a DLN on the Insurance dataset. The model attains a test R2 of 0.866 and relies on two continuous features and two one-hot encoded categorical features, selected from an original set of two continuous and 12 one-hot features.
Figure 8: Decision process learned by a DLN on the Estate dataset. The model attains a test R2 of 0.758 and utilizes all six continuous input features.
Figure 9: Decision process learned by a DLN on the Yacht dataset. The model attains a test R2 of 0.997 and selects three of the six available continuous features.
Orig
No τ sched.
Two phases
Subspace 16
Subspace 4
No concat
Abalone
0.525 ± 0.024
0.524 ± 0.024
0.526 ± 0.027
0.522 ± 0.025
0.520 ± 0.024
0.514 ± 0.015
Airfoil
0.871 ± 0.021
0.830 ± 0.024
0.864 ± 0.016
0.874 ± 0.014
0.866 ± 0.015
0.859 ± 0.018
Bike
0.902 ± 0.016
0.856 ± 0.037
0.900 ± 0.017
0.911 ± 0.016
0.880 ± 0.017
0.854 ± 0.039
CCPP
0.942 ± 2.5e-03
0.939 ± 2.8e-03
0.941 ± 3.1e-03
0.941 ± 3.0e-03
0.940 ± 2.9e-03
0.938 ± 2.9e-03
Concrete
0.887 ± 0.011
0.866 ± 0.023
0.876 ± 0.016
0.882 ± 0.012
0.885 ± 0.016
0.873 ± 0.028
Electrical
0.926 ± 7.2e-03
0.913 ± 0.015
0.932 ± 5.1e-03
0.930 ± 8.6e-03
0.919 ± 8.0e-03
0.918 ± 5.3e-03
Table 7: Average test R2 for the ablation studies across 10 random seeds and 32 hyperparameter trials
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Linear
Ridge
Lasso
KNN
DT
AB
RF
XGB
SVR
MLP
DLN
Abalone
2.16
2.16
2.16
2.20
2.29
2.32
2.13
2.11
2.12
2.07
2.18
Airfoil
4.77
4.78
4.78
2.42
3.10
3.48
2.49
1.50
2.89
1.64
2.28
Bike
140
140
140
106
55.9
105
65.3
37.9
108
44.6
52.5
CCPP
4.54
4.54
4.54
3.52
4.11
4.80
3.90
3.11
3.94
3.86
4.06
Concrete
10.3
10.3
10.3
8.70
7.31
7.38
5.65
4.23
5.84
5.53
5.40
Electrical
0.0220
0.0220
0.0220
0.0166
0.0183
0.0154
0.0147
7.86e-03
7.19e-03
6.57e-03
9.36e-03
Appendix
Table 8: Average test RMSE across 10 random seeds and 128 hyperparameter trials per model
Linear
Ridge
Lasso
KNN
DT
AB
RF
XGB
SVR
MLP
DLN
Abalone
1.57
1.57
1.57
1.54
1.62
1.76
1.50
1.50
1.47
1.48
1.57
Airfoil
3.69
3.69
3.69
1.63
2.34
2.83
1.91
0.993
2.05
1.16
1.71
Bike
105
105
104
69.1
33.4
82.9
41.9
22.9
67.3
27.5
36.5
CCPP
3.61
3.61
3.61
2.48
3.04
3.76
2.95
2.19
2.95
2.87
3.09
Concrete
8.23
8.25
8.27
6.43
5.22
6.01
4.27
2.86
4.29
3.83
3.98
Electrical
0.0174
0.0174
0.0174
0.0129
0.0142
0.0126
0.0116
5.74e-03
4.52e-03
4.38e-03
7.14e-03
Appendix
Table 9: Average test MAE across 10 random seeds and 128 hyperparameter trials per model