Fusion techniques of time frequency-based images to predict the outcome of rTMS depression therapy
Abstract
Depression is a mental condition that can lead to suicide and self-harm. Predicting the outcome of depression treatment is one of the most difficult tasks for clinicians. Among various treatment options, repetitive Transcranial Magnetic Stimulation (rTMS) is a widely used non-invasive method. Predicting rTMS response using Electroencephalogram (EEG) data is difficult because of high inter-subject variability and limited features from single-domain analysis. We introduce two fusion techniques, montage and blending, to overcome these limitations and extract richer features from EEG-derived Time-Frequency (TF) images. We then propose a lightweight custom Convolutional Neural Network (CNN) trained on fused TF representations. \textcolor{black}{We use a primary dataset of 15 patients and a secondary dataset of 46 patients. We run two sets of experiments. The first set uses segment-level 10-fold cross-validation. In this setup segments from the same patient can appear in both training and testing. The Montage CWT_ST fusion reaches 99.90% accuracy on the primary dataset and 91.90% on the secondary dataset. The second set uses strict subject-disjoint cross-validation. All segments of a patient stay in one fold and no patient appears in both training and testing. Performance collapses. We test four time-frequency methods, six fusion mechanisms, and fourteen model architectures. With one exception, every configuration on both cohorts falls between AUC 0.31 and 0.54 and every 95% confidence interval contains 0.5. A patient-level permutation test on the best standalone method returns . The best subject-level result is Montage CWT_ST on the primary cohort, which reaches AUC and 82.7% accuracy.
Figures & tables
| Metric | Responders (R, ) | Non-Responders (NR, ) |
|---|---|---|
| BDI-II Before rTMS (Mean SD) | 27.56 8.19 | 33.83 12.08 |
| BDI-II After rTMS (Mean SD) | 7.44 4.81 | 22.33 10.59 |
| Mean Improvement (%) | 73.20 13.57 | 33.06 11.14 |
| Characteristic | Responders (R, ) | Non-Responders (NR, ) | -Value |
|---|---|---|---|
| Age (Mean SD) | 30.87 12.00 | 39.00 14.16 | 0.052 |
| Gender (Male/Female) | 8/15 | 8/15 | 0.900 |
| BDI Score Before Therapy | 32.5 9.3 | 28.1 9.4 | 0.080 |
| BDI Score After Therapy | 8.6 5.9 | 23.1 8.4 | |
| Length of Depressive Episode (Years) | 6.5 8.2 | 7.9 7.8 | 0.270 |
| Dataset | Precision % | Recall % | Specificity % | % | AUC | Accuracy % | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| NR | R | NR | R | NR | R | NR | R | |||
| DWT | 90 | 90 | 84 | 94 | 90 | 88 | 87 | 92 | 88.99 | 90.22 |
| ST | 98 | 98 | 97 | 99 | 98 | 97 | 97 | 98 | 97.76 | 97.78 |
| SSWE | 98 | 98 | 96 | 99 | 98 | 97 | 97 | 98 | 97.76 | 97.97 |
| CWT | 99 | 99 | 98 | 99 | 99 | 98 | 98 | 99 | 98.52 | 98.67 |
| Model | Precision % | Recall % | F1-score % | Specificity % | AUC | Accuracy % | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| NR | R | NR | R | NR | R | NR | R | |||
| VGG16 | 85 | 83 | 70 | 93 | 77 | 88 | 85 | 83 | 81.24 | 83.94 |
| Xception | 87 | 87 | 77 | 93 | 82 | 90 | 87 | 87 | 85.10 | 87.01 |
| ResNet152V2 | 90 | 87 | 76 | 95 | 83 | 91 | 90 | 87 | 85.60 | 87.86 |
| MobileNetV2 | 91 | 90 | 83 | 95 | 86 | 92 | 91 | 90 | 88.69 | 90.13 |
| DenseNet201 | 91 | 91 | 84 | 95 | 87 | 93 | 93 | 90 | 89.44 | 90.78 |
| Fused | Precision % | Recall % | Specificity % | % | AUC | Accuracy % | ||||
| NR | R | NR | R | NR | R | NR | R | |||
| Montage | ||||||||||
| CWT_DWT | 96 | 97 | 94 | 98 | 96 | 97 | 95 | 97 | 96.05 | 96.48 |
| CWT_SSWE | 98 | 99 | 98 | 99 | 98 | 99 | 98 | 99 | 98.22 | 98.31 |
| CWT_ST | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 99.91 | 99.90 |
| Blend | ||||||||||
| Model | Precision % | Recall % | Specificity % | % | AUC | Accuracy % | ||||
| NR | R | NR | R | NR | R | NR | R | |||
| VGG16 | 98 | 97 | 96 | 99 | 99 | 96 | 97 | 98 | 97.32 | 97.71 |
| Xception | 95 | 94 | 90 | 97 | 97 | 90 | 92 | 96 | 93.55 | 94.40 |
| ResNet152V2 | 97 | 96 | 94 | 98 | 98 | 94 | 95 | 97 | 96.06 | 96.55 |
| MobileNetV2 | 99 | 99 | 98 | 99 | 99 | 98 | 98 | 99 | 98.64 | 98.84 |
| DenseNet201 | 99 | 99 | 98 | 99 | 99 | 98 | 99 | 99 | 98.81 | 98.96 |
| Fused | Precision % | Recall % | Specificity % | % | AUC | Accuracy % | ||||
| NR | R | NR | R | NR | R | NR | R | |||
| Montage | ||||||||||
| CWT_SSWE | 90 | 91 | 90 | 90 | 90 | 90 | 90 | 91 | 90.34 | 90.34 |
| CWT_DWT | 89 | 88 | 89 | 89 | 89 | 88 | 88 | 89 | 88.59 | 88.61 |
| CWT_ST | 92 | 92 | 91 | 92 | 92 | 92 | 92 | 92 | 91.89 | 91.90 |
| Blend | ||||||||||
| Fused | Precision % | Recall % | Specificity % | % | AUC | Accuracy % | ||||
| NR | R | NR | R | NR | R | NR | R | |||
| Montage | ||||||||||
| VGG16 | 72 | 75 | 74 | 74 | 74 | 74 | 73 | 74 | 73.86 | 73.86 |
| ResNet152V2 | 74 | 76 | 76 | 75 | 76 | 75 | 75 | 76 | 75.43 | 75.42 |
| MobileNetV2 | 77 | 78 | 77 | 78 | 77 | 78 | 77 | 78 | 77.72 | 77.74 |
| CNN | 92 | 92 | 91 | 92 | 92 | 92 | 92 | 92 | 91.89 | 91.90 |
| Method | AUC | AUC 95% CI | Sensitivity | Specificity | F1-Score | Accuracy | Brier |
| ST | 0.485 ± 0.079 | 0.170–0.804 | 0.533 ± 0.183 | 0.467 ± 0.217 | 0.552 ± 0.097 | 0.507 ± 0.037 | 0.287 |
| CWT | 0.478 ± 0.072 | 0.138–0.817 | 0.711 ± 0.169 | 0.300 ± 0.183 | 0.647 ± 0.061 | 0.547 ± 0.030 | 0.276 |
| DWT | 0.452 ± 0.085 | 0.135–0.785 | 0.578 ± 0.145 | 0.367 ± 0.183 | 0.569 ± 0.085 | 0.493 ± 0.037 | 0.288 |
| SSWE | 0.441 ± 0.042 | 0.137–0.769 | 0.444 ± 0.208 | 0.567 ± 0.224 | 0.491 ± 0.134 | 0.493 ± 0.037 | 0.292 |
| Strategy | Method | AUC | AUC 95% CI | Sensitivity | Specificity | F1-Score | Accuracy | Brier |
| Blend | CWT_DWT | 0.478 ± 0.063 | 0.147–0.810 | 0.556 ± 0.157 | 0.433 ± 0.149 | 0.565 ± 0.095 | 0.507 ± 0.037 | 0.289 |
| CWT_ST | 0.474 ± 0.098 | 0.147–0.808 | 0.667 ± 0.208 | 0.333 ± 0.204 | 0.617 ± 0.112 | 0.533 ± 0.047 | 0.300 | |
| CWT_SSWE | 0.444 ± 0.113 | 0.139–0.769 | 0.533 ± 0.298 | 0.400 ± 0.303 | 0.521 ± 0.150 | 0.480 ± 0.073 | 0.299 | |
| Montage | CWT_DWT | 0.481 ± 0.090 | 0.158–0.825 | 0.533 ± 0.145 | 0.500 ± 0.167 | 0.563 ± 0.094 | 0.520 ± 0.056 | 0.294 |
| CWT_SSWE | 0.474 ± 0.070 | 0.164–0.792 | 0.400 ± 0.099 | 0.567 ± 0.149 | 0.468 ± 0.058 | 0.467 ± 0.000 | 0.284 | |
| CWT_ST † | 0.874 ± 0.183 | 0.718–0.975 | 0.933 ± 0.149 | 0.667 ± 0.236 | 0.866 ± 0.138 | 0.827 ± 0.174 | 0.130 |
| Model | Family | AUC | Accuracy | Total parameters | Trainable |
| SleepEEGNet | EEG-specific | 0.944 | 0.907 | 771,585 | 771,329 |
| MobileNetV2 | Pretrained | 0.944 | 0.813 | 2,591,297 | 330,753 |
| CNN-LSTM | Hybrid | 0.926 | 0.827 | 974,945 | 974,721 |
| DenseNet201 | Pretrained | 0.922 | 0.827 | 18,821,697 | 495,873 |
| Proposed CNN | Proposed | 0.874 | 0.827 | 407,665 | 407,665 |
| VGG16 | Pretrained | 0.826 | 0.773 | 14,848,321 | 132,609 |
| Model | Family | AUC | Sensitivity | Specificity | F1-Score | Accuracy |
| SVM-RBF | Classical | 0.541 ± 0.079 | 0.578 | 0.667 | 0.639 | 0.613 |
| ResNet152V2 | Pretrained | 0.537 ± 0.112 | 0.778 | 0.233 | 0.660 | 0.560 |
| Logistic regression | Classical | 0.526 ± 0.066 | 0.733 | 0.533 | 0.704 | 0.653 |
| Random forest | Classical | 0.507 ± 0.098 | 0.600 | 0.667 | 0.604 | 0.627 |
| CNN-LSTM | Hybrid | 0.496 ± 0.076 | 0.556 | 0.433 | 0.545 | 0.507 |
| Proposed CNN | Proposed | 0.478 ± 0.072 | 0.711 | 0.300 | 0.647 | 0.547 |
| Model | Family | AUC | Sensitivity | Specificity | F1-Score | Accuracy |
| SleepEEGNet | EEG-specific | 0.944 ± 0.114 | 0.933 | 0.867 | 0.921 | 0.907 |
| MobileNetV2 | Pretrained | 0.944 ± 0.069 | 0.867 | 0.733 | 0.842 | 0.813 |
| CNN-LSTM | Hybrid | 0.926 ± 0.086 | 1.000 | 0.567 | 0.877 | 0.827 |
| DenseNet201 | Pretrained | 0.922 ± 0.082 | 0.822 | 0.833 | 0.841 | 0.827 |
| Proposed CNN | Proposed | 0.874 ± 0.183 | 0.933 | 0.667 | 0.866 | 0.827 |
| VGG16 | Pretrained | 0.826 ± 0.152 | 0.778 | 0.767 | 0.802 | 0.773 |
| Model | Family | AUC | Sensitivity | Specificity | F1-Score | Accuracy |
| Input: standalone ST | ||||||
| SVM-linear | Classical | 0.630 ± 0.081 | 0.609 | 0.722 | 0.620 | 0.665 |
| SVM-RBF | Classical | 0.556 ± 0.069 | 0.722 | 0.513 | 0.641 | 0.617 |
| VGG16 | Pretrained | 0.491 ± 0.091 | 0.426 | 0.565 | 0.407 | 0.496 |
| Random forest | Classical | 0.483 ± 0.047 | 0.817 | 0.348 | 0.662 | 0.583 |
| Proposed CNN | Proposed | 0.470 ± 0.030 | 0.504 | 0.435 | 0.474 | 0.470 |
| Method | AUC | AUC 95% CI | Sensitivity | Specificity | F1-Score | Accuracy | Brier |
| CWT | 0.495 ± 0.033 | 0.320–0.663 | 0.652 ± 0.182 | 0.330 ± 0.136 | 0.554 ± 0.074 | 0.491 ± 0.033 | 0.333 |
| DWT | 0.495 ± 0.017 | 0.321–0.666 | 0.478 ± 0.187 | 0.530 ± 0.200 | 0.478 ± 0.086 | 0.504 ± 0.018 | 0.328 |
| SSWE | 0.479 ± 0.030 | 0.308–0.651 | 0.470 ± 0.094 | 0.530 ± 0.094 | 0.481 ± 0.048 | 0.500 ± 0.000 | 0.347 |
| ST | 0.470 ± 0.030 | 0.299–0.643 | 0.504 ± 0.201 | 0.435 ± 0.215 | 0.474 ± 0.088 | 0.470 ± 0.012 | 0.347 |
| Strategy | Method | AUC | AUC 95% CI | Sensitivity | Specificity | F1-Score | Accuracy | Brier |
| Blend | ST_CWT | 0.529 ± 0.072 | 0.360–0.696 | 0.574 ± 0.235 | 0.426 ± 0.267 | 0.516 ± 0.122 | 0.500 ± 0.074 | 0.334 |
| ST_DWT | 0.505 ± 0.093 | 0.338–0.673 | 0.678 ± 0.245 | 0.278 ± 0.205 | 0.546 ± 0.131 | 0.478 ± 0.027 | 0.332 | |
| ST_SSWE | 0.488 ± 0.064 | 0.321–0.657 | 0.696 ± 0.174 | 0.296 ± 0.223 | 0.573 ± 0.058 | 0.496 ± 0.036 | 0.344 | |
| Montage | ST_CWT | 0.514 ± 0.068 | 0.342–0.686 | 0.617 ± 0.286 | 0.409 ± 0.191 | 0.529 ± 0.176 | 0.513 ± 0.059 | 0.351 |
| ST_SSWE | 0.500 ± 0.089 | 0.331–0.669 | 0.530 ± 0.185 | 0.487 ± 0.200 | 0.505 ± 0.100 | 0.509 ± 0.033 | 0.348 | |
| ST_DWT | 0.457 ± 0.094 | 0.290–0.622 | 0.652 ± 0.378 | 0.339 ± 0.329 | 0.496 ± 0.282 | 0.496 ± 0.054 | 0.342 |
| Quantity | ST | Montage CWT_ST |
|---|---|---|
| Permutations | 100 | 200 |
| Null mean AUC | 0.520 | 0.520 |
| Null standard deviation | 0.083 | 0.105 |
| Null 95th percentile | 0.630 | 0.704 |
| Null maximum | 0.759 | 0.815 |
| Observed AUC, reference seed | 0.500 | 1.000 |
| Experiment | Purpose / Configuration | Filters | Pooling | Parameters | Segment-level (leaky) | Subject-disjoint | ||
| Seg AUC | Patient AUC | Seg AUC | Patient AUC | |||||
| Exp-1 | Baseline shallow CNN | 4,6,8,12 | None | 603,883 | 0.995 | 1.000 | 0.443 | 0.437 |
| Exp-2 | Increased depth, additional 5x5 kernel | 4,6,8,12,32 | None | 1,617,163 | 0.995 | 1.000 | 0.474 | 0.437 |
| Exp-3 | Deepened CNN, larger filter bank | 4,6,8,12,32,64 | None | 3,274,315 | 0.992 | 1.000 | 0.464 | 0.459 |
| Exp-4 | Deep CNN with Average Pooling | 64,128,256 | AvgPool | 205,893,633 | 0.951 | 1.000 | 0.481 | 0.472 |
| Exp-5 | Deep CNN with Global Average Pooling | 64,128,256 | GlobalAvg | 389,121 | 0.861 | 1.000 | 0.456 | 0.461 |
| Fusion mechanism | Where fusion happens | AUC | Sensitivity | Specificity | Accuracy |
| Primary cohort, CWT + ST, stratified group three-fold | |||||
| Montage † | Input, side by side | 0.874 ± 0.183 | 0.933 ± 0.149 | 0.667 ± 0.236 | 0.827 ± 0.174 |
| Blend (alpha overlay) | Input, overlaid | 0.474 ± 0.098 | 0.667 ± 0.208 | 0.333 ± 0.204 | 0.533 ± 0.047 |
| Channel-wise stacking | Input, six channels | 0.504 ± 0.066 | 0.644 ± 0.380 | 0.300 ± 0.415 | 0.507 ± 0.076 |
| Feature concatenation | After two towers | 0.500 ± 0.071 | 0.600 ± 0.169 | 0.367 ± 0.247 | 0.507 ± 0.037 |
| Attention fusion | Weighted feature maps | 0.437 ± 0.084 | 0.778 ± 0.208 | 0.233 ± 0.224 | 0.560 ± 0.037 |
| Ref, Year | Methods | Categorized | Therapy | Number of Patients | Accuracy (%) |
| [ 45 ] , (2023) | Connectivity image of EEG signal channels, ensemble TL model based on voting, LSTM | Time domain | rTMS | 23 R vs. 23 NR | 99.32 |
| [ 40 ] , (2023) | CWT, TL, Bio-LSTM | Time domain | rTMS | 23 R vs. 23 NR | 97.10 |
| [ 44 ] , (2023) | Ensemble TL model, Bio-LSTM | Time domain | rTMS | 23 R vs. 23 NR | 98.51 |
| [ 39 ] , (2023) | Connectivity, ensemble TL models | Time domain | rTMS | 34 depressed patients | 92.28 |
| [ 21 ] , (2018) | Support Vector Machine (SVM) | Time domain | rTMS | 39 patients / 32 patients | 91.00 / 86.66 |
| [ 22 ] , (2023) | Support Vector Machine (SVM) | Time domain | rTMS | 46 R vs. 42 NR | 94.31 |