Differentiable logic gate networks, which operate using only logic gates, have recently attracted attention as an efficient alternative to conventional neural networks. However, despite their efficiency, the scaling behavior of logic gate networks remains underexplored. By contrast, scaling model capacity is a central design principle in deep neural networks and typically leads to improved performance. This discrepancy raises a key question: Can similar scaling benefits also be achieved in logic gate networks? In this work, we focus on width as a primary scaling axis and conduct a systematic analysis of its behavior in logic gate networks. We observe that naive width scaling often introduces redundancy among logic kernels, limiting the effective use of additional kernels and leading to performance saturation. To address this limitation, we propose a dynamic logic kernel framework that reorganizes kernel utilization by promoting specialization across kernel groups. This enables the network to better utilize increased width via input-dependent kernel routing, while ensuring that both routing and computation are implemented entirely with gate-level Boolean operations at inference time. We further find that kernel redundancy is most pronounced at the first gate level, motivating an early-stage dynamic logic kernel strategy that concentrates adaptation at this level. Experimental results demonstrate that our approach improves kernel utilization and increases kernel diversity, leading to higher accuracy with improved parameter efficiency.
Figures & tables
Figure 1: Accuracy under width scaling on the CIFAR-10 dataset.
Figure 2: Function-aware distance between two aligned logic gates.
Method
128
256
384
512
Logic Kernel
3.05
2.50
2.05
1.86
Ours (DLK)
3.05
4.61
4.78
4.89
Table 1: Kernel diversity across widths.
Figure 3: Overview of kernel utilization strategies in logic gate networks: (a) Static utilization, where all kernels are applied without any selection mechanism, (b) dynamic logic kernel (Sec. 4.1 ), (c) early-stage dynamic logic kernel (Sec. 4.2 ).
Figure 4: Level-wise decomposition of kernel diversity.
Figure 5: Top-1 accuracy on CIFAR-10 under width scaling with per-group base width K=128 .
Method
MNIST
SVHN
CIFAR-100
Flowers-102
Tiny-ImageNet
Logic Kernel
98.55
75.58
58.44
32.40
10.55
Logic Kernel+IWP
98.67
76.05
60.33
33.05
10.86
Ours (EDLK)
99.10
77.00
62.46
34.26
13.35
Table 2: Top-1 accuracy across classification benchmarks. The best scores are in bold.
Figure 8Table 9
Method
Latency (ns) ↓
Acc. (%) ↑
Params
Energy (pJ) ↓
Area ( nm2 ) ↓
Logic Kernel
22.06
95.84
21K
4.15
12.55
Logic Kernel
28.43
96.11
42K
5.34
16.15
Ours (EDLK, G=2 )
23.64
97.29
22K
4.37
13.21
Table 6: Hardware cost and classification performance on MNIST.
Method
BOPs ↓
Energy (pJ) ↓
Area (nm 2 ) ↓
Acc. (%) ↑
TTNet [ 29 ]
360K
225
680.75
98.02
DWN [ 30 ]
81K
50.63
153.17
98.20
NeuralLUT [ 31 ]
75K
46.88
141.82
97.72
Ours (EDLK, G=4 )
11K
6.88
20.80
98.25
Table 7: Hardware comparison with recent efficient models on MNIST.
Figure 7: Top-1 accuracy under network depth scaling on CIFAR-10 with fixed width.
Table 8: Representative CNN methods related to adaptive kernel or capacity utilization.
b(g)
Boolean operator g(a1,a2)
00
01
10
11
0000
0
0
0
0
0
0001
a1∧a2
0
0
0
1
0010
a1∧¬a2
0
0
1
0
0011
a1
0
0
1
1
0100
¬a1∧a2
0
1
0
0
0101
a2
0
1
0
1
Appendix
Table 9: Truth-table representation of the 16 two-input Boolean operators.
Figure 8: Per-gate normalized level-wise TED on CIFAR-10. Each level-wise TED value is divided by the number of gates at that level. A similar first-level trend is observed after normalization.