ReFRM3D: A Radiomics-enhanced Fused Residual Multiparametric 3D Network with Multi-Scale Feature Fusion for Glioma Characterization
Authors: Md. Abdur Rahman, Md Noman Hossain, Arefin Ittesafun Abian, Mohaimenul Azam Khan Raiaan, Yan Zhang, Mirjam Jonkman, Sami Azam
Organizations: Applied Artificial Intelligence and Intelligent Systems (AAIINS) Laboratory, Dhaka 1217, Bangladesh · Department of Computer Science and Engineering, United International University, Dhaka, 1212, Bangladesh · Department of Data Science and Artificial Intelligence, Monash University, Clayton, VIC, 3153, Australia · Faculty of Science and Technology, Charles Darwin University, Darwin, 0909, Australia
Gliomas are among the most aggressive cancers, with complex diagnostic processes. Existing glioma segmentation methods often struggle with high variability in imaging data and inadequate optimization. Furthermore, radiomic analysis is typically applied only after segmentation is finished, limiting its ability to inform the segmentation process itself. To address these challenges, we propose a novel radiomics-enhanced fused residual multiparametric 3D network (ReFRM3D) for brain tumor characterization. The framework is based on a 3D U-Net architecture and features multi-scale feature fusion, hybrid upsampling, and an extended residual skip mechanism. Additionally, we introduce a radiomic conditioning mechanism that extracts texture and intensity descriptors from an intermediate coarse segmentation and re-injects them into the decoder to refine the final output. Experimental results on BraTS2019, BraTS2020, and BraTS2021 show strong performance, with mean DSC values of 93.45%, 93.61%, and 92.06%, respectively, across whole tumor, enhancing tumor, and tumor core regions. Compared with recent models, ReFRM3D improves average DSC by 5.79%, 7.25%, and 0.96% on these datasets. Our model also generalized well on the BraTS-Africa dataset with an average DSC of 87.8%, despite differences in scanner field strength and patient demographics.
Figures & tables
Figure 1: Architectural comparison between existing frameworks and the proposed architecture. Top : Standard pipelines utilizing single-scale encoder-decoders, direct skip connections, and basic interpolation. Bottom : The proposed ReFRM3D framework incorporating PCA brain isolation, multi-scale encoding, extended residual skip connections, and a radiomic conditioning loop that extracts descriptors from a coarse mask to guide the final segmentation.
Figure 2: Overview of the proposed glioma segmentation framework. The expansion layers mirror the contraction layers in size and filters.
Figure 3: Visualization of histograms of original and Z-score normalized intensity distributions with intensity comparison across multiple MRI scans. The top-left panel shows the original intensity histogram with high variability (mean μ=270.66 , standard deviation σ=84.43 ), while the top-right panel displays the normalized distribution centered at zero with unit variance. The bottom panel illustrates that while original mean intensities fluctuate significantly across scans (ranging from 290 to 550), the normalized values remain stable (70 to 110), which also ensures consistent input for model training.
Figure 4: Visualization of total slice ranges. Tumors consistently appear between slices 46 and 113 for all 369 volumes (BraTS2020). On average, they occupy 68 slices per volume. The top panel shows tumor slice ranges (gray lines) within the total volume range (light blue bars). Average start and end positions are marked by green and red dashed lines, respectively. The bottom panel displays tumor midpoint positions and shows anatomical consistency across patients, with an average center at slice 79.5. This analysis confirms that excluding non-tumor slices reduces the processed slice count by approximately 56% (from 155 to 68 slices) without affecting segmentation performance.
Figure 5: Overview of the Fused Multi-scale Feature Fusion (FMFF) module. Encoder features from multiple scales undergo channel alignment via 1×1×1 convolutions and spatial alignment via adaptive pooling to form a unified representation.
Figure 6: Overview of the Hybrid Upsampling and Residual Integration (HURI) module. The decoder feature is upsampled using a learnable transposed convolution, integrated with the corresponding encoder feature via a residual connection, and refined through a final convolution block.
Figure 7: Epoch-wise accuracy and loss for training and validation
Figure 8: Qualitative random slice segmentation on a high-grade glioma (HGG) case. Rows show axial, coronal, and sagittal views; columns show T1ce, T2, and FLAIR modalities with overlaid contours, a contour-only ground-truth/prediction comparison, and an overlay. Solid lines denote ground truth and dashed lines denote model predictions, with WT, TC, and ET shown in green, red, and yellow, respectively.
Category
Parameter
Value
Preprocessing
Intensity threshold
99th percentile
Margin cropping ( δ )
1.5 σ
Input size
128×128×128
Input channels
3 (FLAIR, T1ce, T2)
Training Config.
Optimizer
Adam
Learning rate
0.0001
Table 1: Summary of hyperparameters and configuration settings used in ReFRM3D.
Figure 9: Visualization of a random slice of a low-grade glioma (LGG) case. Unlike the HGG case, the LGG lesions are more homogeneous with minimal enhancement.
Modalities
DSC(%)
mmDSC(%)
JCS(%)
mmJCS(%)
WT
ET
TC
WT
ET
TC
BraTS2019
T1ce
92.98
91.17
92.54
92.23
86.88
83.77
85.59
85.41
T2
95.05
94.00
94.71
94.59
90.57
88.68
89.73
89.66
Flair
94.10
92.87
93.66
93.54
88.86
86.69
87.88
87.81
avg.
94.04
92.68
93.64 →
93.45
88.77
86.38
87.73 →
87.63
BraTS2020
T1ce
93.30
90.25
92.74
92.09
87.44
82.23
85.38
85.02
Table 2: Segmentation results for different modalities. The columns mmDSC and mmJCS denote the m ean D ice S imilarity C oefficient and J accard C oefficient S imilarity for the respective m odalities.
Ref
WT
ET
TC
Avg.
Oktay et al. [ 41 ]
0.888
0.759
0.772
0.806
Wang et al. [ 42 ]
0.900
0.789
0.819
0.836
Valanarasu et al. [ 43 ]
0.876
0.732
0.739
0.782
Barzegar et al. [ 18 ]
0.901
0.890
0.887
0.893
Ullah et al. [ 25 ]
0.899
0.845
0.845
0.863
Li et al. [ 27 ]
0.890
0.835
0.842
0.856
Table 3: Comparison of segmentation performance on BraTS2019 dataset. The best and the second-best results are bolded and underlined , respectively.
Ref
WT
ET
TC
Avg.
Oktay et al. [ 41 ]
0.855
0.718
0.759
0.777
Qamar et al. [ 53 ]
0.874
0.794
0.837
0.835
Isensee et al. [ 28 ]
0.850
0.820
0.895
0.855
Wang et al. [ 42 ]
0.900
0.787
0.817
0.835
Xu et al. [ 20 ]
0.900
0.780
0.820
0.833
Jiang et al. [ 44 ]
0.890
0.773
0.803
0.822
Table 4: Comparison of segmentation performance on BraTS2020 dataset. The best and the second-best results are bolded and underlined , respectively.
Ref
WT
ET
TC
Avg.
Hatamizadeh et al. [ 57 ]
0.927
0.868
0.899
0.898
Isensee et al. [ 28 ]
–
–
–
0.912
Jiang et al. [ 44 ]
0.918
0.832
0.847
0.866
Li et al. [ 58 ]
0.923
0.858
0.863
0.881
Lin et al. [ 59 ]
0.933
0.885
0.901
0.906
Roy et al. [ 60 ]
–
–
–
0.915
Table 5: Comparison of segmentation performance on BraTS2021 dataset. The best and the second-best results are bolded and underlined , respectively.
Ref
Dataset
Acc.
Sen
Spec
[ 50 ]
BraTS19
–
75.84%
76.11%
[ 63 ]
BraTS19
98.90%
98.70%
–
[ 52 ]
BraTS19
99.09%
99.15%
99.98%
[ 64 ]
BraTS19-21
98.10%
99.05%
95.20%
[ 63 ]
BraTS20
99.60%
99.00%
–
[ 52 ]
BraTS20
99.88%
99.50%
99.97%
Table 6: Classification performance alongside recent methods on the BraTS benchmark, reported for context only. (–) denotes: values not given.
Model
BraTS
Dice (%)
WT
ET
TC
DDU-net [ 23 ]
2017
0.838
0.783
0.898
DenseNets [ 14 ]
2015
0.816
0.818
0.858
U-Net [ 15 ]
2015
0.733
0.726
0.890
2017
0.763
0.642
0.876
C-ConvNet [ 26 ]
2018
0.872
0.911
0.920
Table 8: Comparison of Dice (%) for different datasets and methods. The best results are highlighted in bold, while the second-best results are underlined.
Method
ET
TC
WT
Avg.
nnU-Net
0.850
0.853
0.913
0.872
MedNeXt Ens.
0.878
0.876
0.924
0.893
BrainUNet
0.759
0.791
0.869
0.806
ReFRM3D (Ours)
0.843
0.880
0.911
0.878
Table 9: Comparison of segmentation performance on the BraTS-Africa dataset.
Metric
Mean ± Std
95% CI
ET
0.843±0.047
[0.832,0.854]
TC
0.880±0.035
[0.872,0.888]
WT
0.911±0.028
[0.905,0.917]
Avg
0.878±0.037
[0.869,0.887]
Table 10: Statistical analysis of ReFRM3D segmentation performance on the BraTS-Africa dataset ( n=146 ). The 95% confidence interval (CI) was estimated from three independent runs to assess result variability.
Figure 10: Visualization of randomly selected slices for three tumor subtypes in the BraTS-Africa dataset.
Ref.
Model
Features
Params (M)
FLOPs (G)
[ 28 ]
nnUNet
CNN
19.07
412.65
[ 57 ]
SwinUNETR
Swin Transformer
61.98
394.84
[ 69 ]
TransUNet
Hybrid CNN‑Transformer
96.07
48.34
[ 70 ]
SegDiff
U‑Net + Diffusion
23
82.1
[ 45 ]
UNETR
Vision Transformer
92.58
41.19
[ 71 ]
MedSegDiff
U‑Net + Diffusion
25
1770
Table 11: Complexity comparison with the SOTA models. Our FLOPs are reported for the input size 128×128×128 .
Metric
∼ Value
Total Params
8.64 M
Trainable Params
8.64 M
Non-trainable Params
4.9 K (BatchNorm stats)
FLOPs ( 1283 input)
263.54 GFLOPs
FLOPs per Param
30.49 K
Mem: weights (FP32, .h5)
68.0 MB
Table 12: System-level complexity of the proposed model.
Component
Parameters
Encoder (L1–L4)
880,288
Bottleneck + Context Pathway
4,426,768
FMFF (1×1 alignment convs)
22,000
Decoder L4 (HURI + rSkip + RadCond.)
2,491,792
Decoder L3 (HURI + rSkip + RadCond.)
623,744
Decoder L2 (HURI + rSkip + RadCond.)
156,352
Table 13: Parameter distribution by architectural component.
Com.
Exp.
Variant
DSC (%)
Mean
Δ Mean
WT
ET
TC
Fusion
F1
No fusion
76.3
72.1
74.8
74.40
0.00
F2
Single-level fusion
78.5
74.3
76.9
76.57
+ 2.17
F3
w/o scale alignment
80.1
76.2
78.7
78.33
+ 3.93
F4
w/o context pathway
81.4
77.8
80.1
79.77
+ 5.37
F5
+ Summation+1 × 1 (FMFF)
82.7
79.3
81.5
81.17
+ 6.77
Table 14: Component-wise ablation study on BraTS2021 (100 randomly selected samples from both HGG and LGG grades, trained for 50 epochs). Each component group is evaluated to isolate its individual contribution. Δ Mean is relative to the baseline of each component group.
Dataset
LR
Loss
Training
Validation
T/E (s)
Accuracy
Loss
Accuracy
Loss
BraTS19
0.01
D/F
0.915
0.232
0.921
0.181
110
0.001
D/F
0.955
0.120
0.969
0.131
119
0.0001
D&F
0.982
0.110
0.970
0.120
125
BraTS2020
0.01
D/F
0.939
0.200
0.939
0.171
108
0.001
D/F
0.960
0.175
0.966
0.130
121
Table 15: Hyperparameter configurations and performance metrics. The columns LR and T/E (s) represent the L earning R ate and the average T ime per E poch in seconds, respectively. This setting uses the Adam optimizer with the input size of 128×128×128 .
Metric
LR
128×128×128
256×256×256
Adam
SGD
Adam
SGD
Acc
0.01
0.959
0.886
0.981
0.917
0.001
0.973
0.929
0.984
0.960
0.0001
0.985
0.941
0.990
0.967
Loss
0.01
0.137
0.156
0.162
0.181
0.001
0.128
0.143
0.136
0.167
Table 16: Controlled optimizer comparison on BraTS2021 (train) with fixed input size and learning rate.
Component
Variant
Tr/Acc
Tr/Loss
Avg DSC (%)
t(s)/Ep
GPU Mem
Inf/T(ms)
Slice Selection
Full range ∗
0.891
0.27
81.34
151±2.1
2,847 MB
312±8
Tumor-only
0.929
0.21
83.89
123±2.4
2,413 MB
264±5
Normalization
No normalization ∗
0.891
0.27
81.34
151±2.1
2,847 MB
312±8
Min-max
0.911
0.23
82.54
151±1.9
2,847 MB
312±7
Z-score
0.929
0.20
83.12
151±1.7
2,847 MB
312±6
Brain Isolation
No isolation ∗
0.891
0.27
81.34
151±2.1
2,847 MB
312±8
Table 17: Preprocessing component ablation. GPU memory and inference time are reported for a single volume forward pass. Training time (per epoch) and inference time are reported in the t(s)/Ep and Inf/T(ms) columns, respectively. Baselines are indicated with ” ∗ ”.
Cropping
Acc
Loss
DSC (%)
t(s)/Ep
Axis-aligned
0.931
0.28
84.21
141
PCA (0.75 σ )
0.938
0.23
84.89
108
PCA (1.00 σ )
0.943
0.19
85.34
112
PCA (1.25 σ )
0.944
0.17
85.73
119
PCA (1.50 σ )
0.944
0.17
85.73
123
PCA (1.75 σ )
0.941
0.21
80.89
129
Table 18: Sensitivity analysis of brain region cropping strategy and margin parameter on segmentation performance and training efficiency. Training time (per epoch) is reported in the t(s)/Ep column.
Department of Computer Science, Dr. Bhimrao Ambedkar University, Agra, India. · Department of Computer Science and Engineering, Shiv Nadar University, Greater Noida, India.