U-Net inference for brain-tumor segmentation requires billions of multiply-accumulate operations, motivating hardware that can reduce computation dynamically rather than relying only on fixed precision or static model compression. Most-significant-digit-first (MSDF) arithmetic exposes the leading digits of a result during computation, enabling output-dependent decisions before the full value is generated. This paper presents an MSDF accelerator for quantized U-Net segmentation with a two-stage grouped processing element supporting signed INT8 operands and in-stream bias accumulation. Four runtime mechanisms operate directly on the output digit stream: exact early negative detection (END) in ReLU layers, exact sign-only decision making in the segmentation head, calibrated low-order-digit skipping, and calibrated pruning. The two approximate mechanisms are selected offline under an accuracy constraint, while execution requires only lightweight control and does not modify the stored weights. On a residual U-Net trained with nnU-Net for BraTS, the proposed mechanisms reduce digit cycles by 38.38% while achieving a mean Dice score of 80.58% on 73 held-out cases, compared with 81.20% for the floating-point model; the exact mechanisms alone reduce cycles by 18.79% without altering the quantized output. Synthesized in 45nm, the processing element operates at 500MHz, occupies 0.858mm2, and consumes 0.726mJ per 192×192 patch under switching-activity-annotated power analysis. A projected eight-output accelerator with shared activation delivery achieves 16.6ms latency and 1.67mJ per patch.
Figures & tables
Figure 1: Evaluated residual U-Net, drawn by spatial resolution: the encoder descends on the left, the bottleneck (same resolution as e3 ) sits at the bottom, and the decoder ascends on the right. Each box lists its two convolutions (kernel, channels) and its output tensor; the violet tag marks the parameter-free residual connection of the stage; the yellow tags give the number of Stage-1 groups of the processing element that each convolution enables (Section 5 ). Each concatenation node receives the encoder feature at the lower resolution, and the concatenated tensor is upsampled by two (orange) before the next pair of convolutions.
Figure 2: Processing element architecture. (a) Single-stage MMA organization from [ 9 ] , where the reduction depth increases with the number of lanes. (b) Proposed two-stage grouped MSDF PE with local group reductions, second-stage accumulation, and OGF-based signed-digit generation.
Figure 3: Lane datapath, operand stream, and runtime control of the PE. (a) Signed partial-product selection for each activation plane. (b) Convolution transaction including N=KhKwCin products and bias pseudo-products. (c) Priority controller for early termination.
Layer
Input ( C×H×W )
Output
Kernel
Cin
Cout
N
M
G / G2
Packing (%)
Transactions
Lfull
conv0_0
4×1922
16×962
3×3
4
16
36
64
1 / 1
58.8
147,456
24
conv0_1
16×962
16×962
3×3
16
16
144
32
5 / 8
91.3
147,456
26
conv1_0
16×962
32×482
3×3
16
32
144
32
5 / 8
91.2
73,728
26
conv1_1
32×482
32×482
3×3
32
32
288
64
5 / 8
90.6
73,728
27
conv2_0
32×482
64×242
3×3
32
64
288
64
5 / 8
90.6
36,864
27
conv2_1
64×242
64×242
3×3
64
64
576
64
10 / 16
90.3
36,864
28
Table 1: Exact Mapping of the Quantized Network onto the 64-Group PE ( B=8 )
Figure 4: Early-termination policies operating on the MSDF digit stream. (1) Exact early negative detection (END) of ReLU pre-activations. (2) Calibrated pruning of near-zero outputs. (3) Calibrated skipping of low-order digits. (4) Exact sign-based decision making in the segmentation head.
s / p
Mean
WT
TC
ET
Cyc. saved
FP32 ref.
81.20
90.45
76.04
77.10
–
off / off
81.00
90.31
76.00
76.68
18.79%
skip calculation only
6 / off
80.74
90.28
76.12
75.82
33.44%
9 / off
79.93
88.01
75.69
76.10
39.47%
10 / off
72.83
75.81
71.20
71.48
42.04%
Table 2: Accuracy and Digit-Cycle Reduction of the Calibration Sweep
Figure 5: Calibration sweep of the skip exponent s and the pruning exponent p . (a) Mean Dice versus cycles saved for all 49 configurations; each curve steps through the pruning settings of one skip exponent. (b) Zoom on the operating region with the ±2 -point gate; the star marks the selected s=8 , p=9 .
Figure 6: Segmentation outputs of two held-out cases (axial T1-gd slices) for the ground truth, the floating-point model, the INT8 model with exact policies only, the selected operating point ( s=8 , p=9 , shaded), and three configurations beyond the accuracy constraint. Numbers under each panel are the case-level WT / TC / ET Dice scores in %; numbers under the column headings are the means over the 73 evaluation cases.
Mode
Groups
Lanes
Power (mW)
pJ/cycle
MG1-64L
1
64
7.48
14.97
MG5-32L
5
32
7.36
14.73
MG5-64L
5
64
7.62
15.24
MG10-64L
10
64
7.80
15.60
MG19-64L
19
64
8.14
16.29
MG37-64L
37
64
8.87
17.73
Table 3: SAIF-Annotated PE Modes at 500 MHz (Nangate 45 nm, Typical Corner)
Parameter
Meaning
Value
Source/Basis
APE
PE cell area
0.858 mm 2
Synth.
EPE
PE energy/patch
0.7262 mJ
SAIF
Cmeas
Cycles/patch
45.712 M
Sim.
f
Clock frequency
500 MHz
Timing
MSRAM
SRAM capacity
3.289 MiB
Buffer
Brd/wr/ext
Traffic volumes
30.7/5.1/3.1 MiB
Schedule
Table 4: Parameters Used in the System-Level Projection
Figure 7: Scaling behavior of the projected output-tiled accelerator. (a) Latency decomposition showing the reduction of convolution digit-stream execution and activation delivery overhead with increasing output tile size. (b) Effective parallelism and activation reuse achieved by output tiling. (c) Energy breakdown showing that reduced execution time and memory movement dominate the additional compute resources.
Output lanes To
1
2
4
8
Peff
1.00
1.99
3.97
7.70
Activation reuse
1.00 ×
2.00 ×
3.99 ×
7.91 ×
Area (mm 2 )
13.0
14.1
16.3
20.9
Latency (ms)
99.8
52.3
28.2
16.6
Patches/s
10.0
19.1
35.5
60.4
Energy (mJ)
3.30
2.32
1.85
1.67
Table 5: Projected Output-Tile Sweep at the Selected Operating Point
Work
Computation style
Runtime reduction
Technology
Application
Reported result
Stripes [ 15 ]
Bit-serial precision scaling
Compile-time
n/r
CNN classification
Precision adaptation
UNPU [ 18 ]
Bit-serial variable precision
Compile-time
65 nm
CNN classification
Silicon prototype
BitSET [ 40 ]
Bit-serial partial-sum prediction
Runtime prediction
45 nm
CNN classification
1.5 × speedup; 1.4 × energy improvement
BitFair [ 41 ]
Learned bit ordering and thresholding
Runtime prediction
12 nm FinFET
Event-based recognition
Up to 234 BTOPS/W
DSLR-CNN [ 14 ]
MSDF digit-serial computation
None
45 nm
CNN classification
3.57 TOPS/W peak
On-CNN [ 19 ]
MSDF online arithmetic
Runtime power mode
n/r
CNN classification
Up to 33.8% power reduction
Table 6: Comparison of Digit- and Bit-Serial Accelerators With Runtime Computation Reduction
Work / task
Platform / implementation
Power
Latency (ms)
Energy/frame (mJ)
GOPS/W
Accuracy
Liu et al. [ 24 ] / Cityscapes 5122
Zynq ZC706, 16-bit fixed; board measurement
9.60 W
58.8
564.7
11.1
60.8% pixel acc.
Zheng et al. [ 27 ] / DRIVE vessel segmentation
VCU128, INT16; FPGA implementation
16.708 W
26.3
439.4
67.4
64.97% mIoU
Sang et al. [ 29 ] / 2-D U-Net
CGLA, 28-nm synthesis projection
2.61 W
121
315.8
18.68
n/r
Wang [ 28 ] / Medical segmentation 2562
Zynq-7000, INT8; FPGA evaluation
9.587 W
6.8
65.2
7.3
n/r
Usman et al. [ 9 ] / U-Net convolution layers
Zynq-7020, INT8, 100 MHz; board measurement
3.50 W
53.25
186.2
15.14
n/r
Xiong et al. [ 42 ] / BraTS tumor segmentation
Alveo U280, INT8; board measurement
45 W
150
6750
n/r
0.871/0.882 DSC
Table 7: Comparison With U-Net and Segmentation Hardware Accelerators. Reported energy and efficiency values are calculated from the reported power and throughput when not directly provided.
This paper presents an energy-efficient hardware acceleration of the convolutional layers in the U-Net architecture for image segmentation, implemented on FPGA. While digit-serial arithmetic, particularly most-significant-digit-first (MSDF) techniques, offers a compact hardware footprint, it suffers from initial latency before producing the first output digit. This delay accumulates in cascaded operations like multiplication followed by addition, where each unit introduces its own startup overhead. To overcome this, we propose a merged multiply-add (MMA) architecture that fuses these operations into a unified pipeline. Instead of incurring separate delays, the MMA introduces a single streamlined latency per iteration, shorter than the combined latency of conventional cascaded units, resulting in enhanced throughput and efficiency. The MMA units are designed to process spatial input depths in parallel, achieving significantly higher performance than both standalone MSDF-based and conventional designs. We evaluate the proposed design using U-Net as a target application. Despite operating at a lower frequency than a CPU, the FPGA-based accelerator achieves up to an order of magnitude higher energy efficiency, delivering up to 15.14 GOPS/W compared to 1.93 GOPS/W for CPU-based inference. The design also shows approximately 9× reduction in energy consumption compared to MSDF-based FPGA implementations. These results highlight the efficacy of the merged arithmetic approach for resource-constrained, latency-sensitive edge applications in medical imaging and computer vision.
Muhammad Usman, Yousef Sadegheih, Dorit Merhof
Faculty of Informatics and Data Science, University of Regensburg, 93053 Regensburg, Germany
Automatic brain tumor segmentation from multi-modal MRI remains challenging because volumetric models often incur substantial computational cost. This paper presents DALight-3D, a compact 3D U-Net variant that combines depthwise separable 3D convolutions, identifier-conditioned normalization, cross-slice attention, and adaptive skip fusion. The method is evaluated on the Medical Segmentation Decathlon Task01 BrainTumour benchmark under matched optimization settings against standard 3D U-Net, Attention U-Net, Residual 3D U-Net, and V-Net baselines. In the reported 50-epoch comparison, DALight-3D achieves a mean Dice of 0.727 with 2.22M parameters, compared with 0.710 Dice and 3.20M parameters for Residual 3D U-Net. Component-wise ablations show consistent performance degradation when SepConv, identifier-conditioned normalization, CSA, or SSFB is removed. These results indicate that DALight-3D offers a favorable accuracy-efficiency trade-off within the present benchmark setting.
Nand Kumar Mishra, Dhruv Mishra, Dr Manu Pratap Singh
Department of Computer Science, Dr. Bhimrao Ambedkar University, Agra, India. · Department of Computer Science and Engineering, Shiv Nadar University, Greater Noida, India.
Accurate 3D brain tumour segmentation from multi-modal Magnetic Resonance Imaging (MRI) is essential for clinical diagnosis and treatment planning. Existing brain tumour segmentation methods often suffer from heavy computational demands, while current lightweight architectures frequently lack the capacity to maintain segmentation fidelity in complex tumour regions. To address these issues, we propose a novel ultra-lightweight framework (Uni-Light) that achieves high-fidelity segmentation with substantially reduced computational overhead. It combines multi-scale convolutions with an uncertainty-aware knowledge distillation scheme that directs the student model toward hard-to-classify regions, complemented by a Signed Distance Field boundary loss for geometric constraints. Experimental results on BraTS2023-GLI and MSD-BTS datasets demonstrate that Uni-Light reduces parameters by 97.56%, floating-point operations (FLOPs) by 73.03%, and inference memory footprint by 81.58%, while surpassing the state-of-the-art model by an average of 1.47% in Dice score, offering a highly competitive trade-off between segmentation accuracy and computational efficiency in resource-constrained clinical settings. This work also advances data engineering for medical imaging by demonstrating that teacher model uncertainty can be exploited as a data-driven supervisory signal, re-prioritising the training data distribution without requiring additional annotation.
Libing Kuang, Soren Salehi, Ziling Wu +2
School of Computer Science, University of Nottingham, Nottingham, UK · Electrical Engineering Department, Sharif University of Technology, Tehran, Iran · University of Pittsburgh, Pittsburgh, PA, USA