UDSS-BWE: Uncertainty- and Decision-Science Inspired Swin BandWidth Extension
Organizations: Department of Cyber Security Engineering, George Mason University, USA.
Abstract
Bandwidth extension (BWE) is fundamentally localized: the most perceptual distortions are not average-case distortions, but rare high-frequency (HF) transients that standard, risk-neutral objectives tend to smooth away. To close this gap, we seek solutions in the risk-sensitive and uncertainty-aware decision science rules and present UDSS-BWE, which introduces five decision-science and uncertainty-aware discriminators: CVaRD (does tail pooling to amplify HF artifacts), CCD (a primal-dual augmented Lagrangian to prevent HF overboost), MCUD (a learnable utility over spectral flatness/ centroid/ rolloff), EDD (captures epistemic uncertainty), and DROD (captures entropic KL-DRO aggregation). UDSS-BWE is also designed as a complex valued adversarial BWE framework that uses Swin-based generators, a lightweight dual-stream shifted-window backbone, to capture local and long-range structure efficiently, while learnable lattice coupling provides controlled cross-stream exchange. UDSS-BWE is optimized extensively and achieves better perceptual quality with 3.89x fewer parameters (72M vs.18.5M) over two English and French datasets under clean and noisy conditions. To the best of our knowledge, this work shows how multi disciplinary decision-science-inspired and uncertainty theories can be successfully used to design efficient discriminators for producing more nuanced audios, establishing a new baseline in the BWE task.
Figures & tables
| Method | Size | Data | NISQA-MOS | STOI | PESQ | SI-SDR | SI-SNR | LSD | WER % | ||||||||||||||
| 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | |||
| Unprocessed | - | VCTK | 2.79 | 3.67 | 4.43 | 0.55 | 0.61 | 0.61 | 1.15 | 1.51 | 1.41 | -11.03 | -8.07 | -6.07 | -10.53 | -7.62 | -5.63 | 3.27 | 2.27 | 2.85 | 90.0 | 6.1 | 2.1 |
| MLS | 2.29 | 3.12 | - | 0.51 | 0.57 | - | 1.05 | 1.34 | - | -14.15 | -18.89 | - | -13.67 | -17.71 | - | 3.48 | 2.48 | - | 91.1 | 6.8 | - | ||
| EBEN, 2023 | 29.7M | VCTK | 2.59 | 2.69 | 2.53 | 0.89 | 0.98 | 0.98 | 2.64 | 3.69 | 3.71 | 11.94 | 19.94 | 20.82 | 11.94 | 19.94 | 20.83 | 1.03 | 0.78 | 0.92 | 19.1 | 12.4 | 8.5 |
| MLS | 2.14 | 2.28 | - | 0.87 | 0.96 | - | 2.47 | 3.55 | - | 11.81 | 18.24 | - | 11.74 | 18.38 | - | 1.17 | 0.88 | - | 19.4 | 12.7 | - | ||
| AERO, 2023 | 36.4M | VCTK | 2.79 | 2.75 | 2.88 | 0.83 | 0.94 | 0.99 | 2.62 | 3.65 | 3.69 | 13.60 | 20.70 | 21.56 | 13.60 | 20.70 | 21.56 | 1.09 | 0.97 | 0.75 | 20.3 | 13.4 | 9.8 |
| SL | A | P | CVaR | CC | MCU | ED | DRO | L | S | P | SN | N |
| Single discriminators (no MRAD/MRPD) | ||||||||||||
| 1 | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | 1.11 | 0.95 | 2.53 | 14.31 | 3.07 |
| 2 | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | 1.14 | 0.95 | 2.65 | 14.46 | 3.05 |
| 3 | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | 1.99 | 0.95 | 2.62 | 14.16 | 3.02 |
| 4 | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | 1.20 | 0.92 | 1.97 | 10.22 | 2.98 |
| 5 | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | 1.13 | 0.95 | 2.54 | 14.42 | 3.03 |
| Method | Data | NISQA-MOS | SI-SNR | LSD | ||||||
| 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | ||
| EBEN | VCTK | 1.01 | 1.08 | 1.15 | 4.23 | 5.31 | 6.01 | 1.41 | 1.11 | 1.00 |
| AERO | VCTK | 1.52 | 1.12 | 1.19 | 4.39 | 5.38 | 6.15 | 1.43 | 1.13 | 1.01 |
| AP-BWE | VCTK | 2.74 | 3.15 | 3.71 | 4.36 | 5.31 | 6.11 | 1.37 | 1.07 | 0.89 |
| UDSS-BWE | VCTK | 3.91 | 3.87 | 3.94 | 4.01 | 5.65 | 7.01 | 1.22 | 1.02 | 0.89 |
| Condition | LSD | SI-SNR | N-MOS |
| Ideal/PCM-24 | 0.95 (R0) | 12.98 (R0) | 4.42 (R3) |
| PCM-8 | 1.00 (R1) | 12.37 (R0) | 3.76 (R1) |
| G.711 -law | 0.95 (R0) | 12.93 (R0) | 4.40 (R3) |
| G.711 A-law | 0.95 (R0) | 12.95 (R0) | 4.38 (R3) |
| Loss 3/5% | 1.10/1.10 (R2) | 7.79/6.93 (R4) | 3.53/3.30 (R1) |
| Model | Freq. | Par.(M) | MAC(M) | FLOP(M) | RTF(GPU) | Inf.(ms) |
| AP-BWE | 16-48 kHz | 72.07 | 14236.65 | 28473.31 | 0.0025x | 16.60 |
| UDSS-BWE | 16-48 kHz | 18.5 | 5334.72 | 10669.45 | 0.0028x | 16.94 |
| Row | Band | ||||||
| AP | UDSS | AP | UDSS | AP | UDSS | ||
| \scriptsize1⃝ | 0–4 | 0.0521 | 0.0514 | 0.227 | 0.302 | ||
| \scriptsize2⃝ | 4–8 | 0.375 | 0.171 | 0.704 | 0.561 | ||
| \scriptsize3⃝ | 8–12 | 0.00114 | 0.000906 | 0.831 | 0.774 | 3.979 | 4.233 |
| \scriptsize4⃝ | 12–16 | 0.00124 | 0.000866 | 0.689 | 1.018 | 3.214 | 2.961 |
| \scriptsize5⃝ | 16–20 | 0.00113 | 0.000795 | 0.775 | 0.793 | 3.616 | 2.560 |
| Agg. | Method | Mean | SD | SEM | 95% CI | |
| Utt. | AP-BWE | 1000 | 4.458 | 0.472 | 0.015 | [4.429, 4.488] |
| Utt. | UDSS-BWE | 1000 | 4.497 | 0.486 | 0.015 | [4.466, 4.527] |
| Spk. | AP-BWE | 7 | 4.448 | 0.196 | 0.074 | [4.267, 4.630] |
| Spk. | UDSS-BWE | 7 | 4.487 | 0.196 | 0.074 | [4.305, 4.668] |
| Agg. | Median | ||||||
| Utt. | 1000 | 0.038 | 0.038 | 363749 | 136751 | 136751 | |
| Spk. | 7 | 0.038 | 0.039 | 27 | 1 | 1 | 0.0313 |
| Feature Extractor (FE) | MACs (M) | FLOPs (M) | Latency (ms) |
| CVaR_FE | 49.696 | 99.392 | 0.659 |
| CC_FE | 0.141 | 0.282 | 0.823 |
| MCU_FE | 0.139 | 0.278 | 0.889 |
| ED_FE | 14.376 | 28.752 | 0.441 |
| DRO_FE | 14.248 | 28.496 | 0.348 |
Appendix figures & tables36 assets
Supplementary material from the paper’s appendix.
Appendix
| Loss function | Equation (MSE = Mean Squared Error) | Terms |
| Feature Matching Loss | : layer- feature map from ; : per-d FM weight. | |
| Generator Hinge Loss | ||
| (base adversarial set) | = | ; : scalar hinge critic score of discriminator . |
| Discriminator Hinge Loss | ||
| (base adversarial set) | = | For each : real-hinge enforces and fake-hinge enforces . |
| EDD Discriminator Loss | : Dirichlet concentration over ; , ; . |
| Condition | Overall | 8-kHz | 16-kHz |
| Soft speech | 54.8 | 41.8 | 32.2 |
| Normal speech | 62.0 | 46.4 | 38.0 |
| Loud speech | 73.8 | 54.7 | 47.6 |
| Normal singing | 73.9 | 50.2 | 42.3 |
| Fricative | 8-kHz | 16-kHz |
| /s/ | 57.0 | 48.2 |
| /sh/ | 54.9 | 37.7 |
| /f/ | 39.2 | 38.7 |
| /th/ | 36.7 | 39.5 |
| Model | Range | WER | CER | Word Accuracy |
| Unprocessed | 4 kHz | 90.01% | 68.59% | 10.12% |
| UDSS-BWE | 4-16 kHz | 13.6 % | 10.19% | 82.74% |
| Unprocessed | 8 kHz | 6.14% | 3.15% | 93.96% |
| UDSS-BWE | 8-16 kHz | 4.18% | 2.69% | 94.87% |
| Unprocessed | 16 kHz | 2.1% | 1.21% | 97.47% |
| UDSS-BWE | 16-48 kHz | 1.61% | 0.79% | 98.4% |
| Model | SNR | Samples | NISQA-MOS | STOI | PESQ | SI-SDR | SI-SNR | LSD | ||||||||||||||
| 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | ||
| AP-BWE | -10 | 558 | 558 | 593 | 2.02 | 2.55 | 2.98 | 0.60 | 0.72 | 0.73 | 1.08 | 1.19 | 1.15 | -0.52 | 0.05 | 1.24 | -0.60 | -0.04 | 1.17 | 1.49 | 1.18 | 0.91 |
| -5 | 585 | 585 | 608 | 2.52 | 2.86 | 3.54 | 0.69 | 0.80 | 0.80 | 1.15 | 1.38 | 1.26 | 2.96 | 3.62 | 4.54 | 2.86 | 3.48 | 4.44 | 1.45 | 1.11 | 0.88 | |
| 0 | 606 | 606 | 601 | 2.83 | 3.21 | 3.80 | 0.76 | 0.86 | 0.85 | 1.24 | 1.63 | 1.41 | 5.13 | 6.03 | 6.87 | 5.06 | 5.93 | 6.80 | 1.39 | 1.06 | 0.86 | |
| 5 | 580 | 580 | 576 | 3.05 | 3.42 | 4.07 | 0.81 | 0.90 | 0.89 | 1.35 | 1.93 | 1.61 | 6.57 | 7.81 | 8.68 | 6.53 | 7.74 | 8.64 | 1.34 | 1.02 | 0.84 | |
| 10 | 608 | 608 | 559 | 3.22 | 3.70 | 4.22 | 0.84 | 0.93 | 0.92 | 1.49 | 2.26 | 1.80 | 7.63 | 9.05 | 9.77 | 7.62 | 9.02 | 9.77 | 1.29 | 0.97 | 0.82 | |
| Model | Noise Type | Samples | NISQA-MOS | STOI | PESQ | SI-SDR | SI-SNR | LSD | ||||||||||||||
| 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | 4–16 | 8–16 | 16–48 | ||
| AP-BWE | AWGN | 348 | 348 | 378 | 2.68 | 2.73 | 3.76 | 0.75 | 0.81 | 0.80 | 1.25 | 1.52 | 1.35 | 4.99 | 6.33 | 6.54 | 4.90 | 6.22 | 6.52 | 1.40 | 2.73 | 0.88 |
| Airport | 372 | 372 | 411 | 2.73 | 3.27 | 3.89 | 0.74 | 0.86 | 0.85 | 1.28 | 1.75 | 1.49 | 4.44 | 5.34 | 6.28 | 4.41 | 5.28 | 6.22 | 1.39 | 3.27 | 0.83 | |
| Babble | 384 | 384 | 348 | 2.67 | 3.26 | 3.82 | 0.74 | 0.85 | 0.84 | 1.28 | 1.70 | 1.45 | 4.28 | 5.01 | 5.93 | 4.23 | 4.94 | 5.88 | 1.38 | 3.26 | 0.83 | |
| Car | 354 | 354 | 370 | 2.87 | 3.23 | 3.66 | 0.75 | 0.85 | 0.84 | 1.26 | 1.71 | 1.45 | 4.55 | 5.42 | 6.26 | 4.49 | 5.34 | 6.21 | 1.38 | 3.23 | 0.85 | |
| Exhibition | 333 | 333 | 366 | 2.91 | 3.24 | 3.94 | 0.77 | 0.86 | 0.85 | 1.30 | 1.79 | 1.47 | 5.41 | 6.15 | 6.91 | 5.37 | 6.07 | 6.87 | 1.34 | 3.24 | 0.84 | |
| Feature Extractor | MACs (M) | FLOPs (M) | Latency (ms) |
| CVaR_FE | 49.696 | 99.392 | 0.659 |
| CC_FE | 0.141 | 0.282 | 0.823 |
| MCU_FE | 0.139 | 0.278 | 0.889 |
| ED_FE | 14.376 | 28.752 | 0.441 |
| DRO_FE | 14.248 | 28.496 | 0.348 |
| Freq. range | LSD | STOI | PESQ | SDR | SNR | N-MOS | WER |
| 2-16 kHz | 1.08 | 0.86 | 1.5 | 7.61 | 7.59 | 4.49 | 0.60 |
| 2-48 kHz | 1.07 | 0.84 | 1.42 | 7.24 | 7.25 | 4.04 | 0.73 |
| 4-48 kHz | 1.28 | 0.89 | 1.37 | 9.23 | 9.32 | 2.57 | 0.35 |
| 8-48 kHz | 0.94 | 0.99 | 3.37 | 15.32 | 15.24 | 4.52 | 0.05 |
| 12-48 kHz | 0.87 | 0.99 | 4.23 | 16.81 | 16.69 | 4.54 | 0.03 |
| 24-48 kHz | 0.67 | 0.99 | 4.47 | 22.23 | 22.21 | 4.52 | 0.01 |
| Row | Band | ||||||
| AP | UDSS | AP | UDSS | AP | UDSS | ||
| \scriptsize1⃝ | 0–4 | 0.0521 | 0.0514 | 0.227 | 0.302 | ||
| \scriptsize2⃝ | 4–8 | 0.375 | 0.171 | 0.704 | 0.561 | ||
| \scriptsize3⃝ | 8–12 | 0.00114 | 0.000906 | 0.831 | 0.774 | 3.979 | 4.233 |
| \scriptsize4⃝ | 12–16 | 0.00124 | 0.000866 | 0.689 | 1.018 | 3.214 | 2.961 |
| \scriptsize5⃝ | 16–20 | 0.00113 | 0.000795 | 0.775 | 0.793 | 3.616 | 2.560 |
| Agg. | Method | MOS | SD | SEM | 95% CI | |
| Utt. | AP-BWE | 1000 | 4.458 | 0.472 | 0.015 | [4.429, 4.488] |
| Utt. | UDSS-BWE | 1000 | 4.497 | 0.486 | 0.015 | [4.466, 4.527] |
| Spk. | AP-BWE | 7 | 4.448 | 0.196 | 0.074 | [4.267, 4.630] |
| Spk. | UDSS-BWE | 7 | 4.487 | 0.196 | 0.074 | [4.305, 4.668] |
| Agg. | Median | ||||||
| Utt. | 1000 | 0.038 | 0.038 | 363749 | 136751 | 136751 | |
| Spk. | 7 | 0.038 | 0.039 | 27 | 1 | 1 | 0.0313 |
| Row | Test condition | Exp. | LSD | STOI | PESQ | SI-SDR | SI-SNR | N-MOS |
| \scriptsize1⃝ | Ideal | R0 | 0.9471 | 0.9445 | 2.4250 | 13.0330 | 12.9754 | 4.2779 |
| \scriptsize2⃝ | R1 | 0.9533 | 0.9412 | 2.3440 | 12.4628 | 12.3991 | 4.4166 | |
| \scriptsize3⃝ | R2 | 0.9514 | 0.9458 | 2.3629 | 12.4504 | 12.3844 | 4.3511 | |
| \scriptsize4⃝ | R3 | 0.9586 | 0.9397 | 2.2754 | 12.4433 | 12.3627 | 4.4190 | |
| \scriptsize5⃝ | R4 | 1.2455 | 0.8781 | 1.5612 | 2.6345 | |||
| \scriptsize6⃝ | PCM-24 | R0 | 0.9471 | 0.9446 | 2.4249 | 13.0326 | 12.9751 | 4.2783 |
| Row | Test condition | Exp. | LSD | STOI | PESQ | SI-SDR | SI-SNR | N-MOS |
| \scriptsize1⃝ | Loss-3 | R0 | 1.1206 | 0.8921 | 1.7298 | 2.9929 | ||
| \scriptsize2⃝ | R1 | 1.1186 | 0.8864 | 1.6759 | 3.5332 | |||
| \scriptsize3⃝ | R2 | 1.1012 | 0.8962 | 1.7421 | 3.3267 | |||
| \scriptsize4⃝ | R3 | 1.1426 | 0.8868 | 1.5979 | 3.4352 | |||
| \scriptsize5⃝ | R4 | 1.1346 | 0.9104 | 1.8614 | 7.8475 | 7.7938 | 2.9225 | |
| \scriptsize6⃝ | Loss-5 | R0 | 1.1209 | 0.8824 | 1.6314 | 2.7699 |
| Discriminator | Stage | Layer Type | In→Out | Kernel | Stride | Padding | Params |
| CVaRD | Block 1 | Depthwise Conv1d | 1→1 | 7 | 2 | 3 | 8 |
| Pointwise Conv1d | 1→32 | 1 | 1 | 0 | 64 | ||
| BatchNorm1d + LReLU(0.2) | 32→32 | – | – | – | 64 | ||
| Block 2 | Depthwise Conv1d | 32→32 | 7 | 2 | 3 | 256 | |
| Pointwise Conv1d | 32→64 | 1 | 1 | 0 | 2 112 | ||
| BatchNorm1d + LReLU(0.2) | 64→64 | – | – | – | 128 |
| Discriminator | Stage | Layer Type | In→Out | Kernel | Stride | Padding | Params |
| CCD | Block 1 | Depthwise Conv1d | 1→1 | 9 | 2 | 4 | 10 |
| Pointwise Conv1d | 1→32 | 1 | 1 | 0 | 64 | ||
| BatchNorm1d + LReLU(0.2) | 32→32 | – | – | – | 64 | ||
| Block 2 | Depthwise Conv1d | 32→32 | 7 | 2 | 3 | 256 | |
| Pointwise Conv1d | 32→64 | 1 | 1 | 0 | 2 112 | ||
| BatchNorm1d + LReLU(0.2) | 64→64 | – | – | – | 128 |
| Discriminator | Stage | Layer Type | In→Out | Kernel | Stride | Padding | Params |
| EDD | Block 1 | Depthwise Conv1d | 1→1 | 9 | 2 | 4 | 10 |
| Pointwise Conv1d | 1→32 | 1 | 1 | 0 | 64 | ||
| Block 2 | Depthwise Conv1d | 32→32 | 7 | 2 | 3 | 256 | |
| Pointwise Conv1d | 32→64 | 1 | 1 | 0 | 2 112 | ||
| Block 3 | Depthwise Conv1d | 64→64 | 5 | 2 | 2 | 384 | |
| Pointwise Conv1d | 64→64 | 1 | 1 | 0 | 4 160 |
| Discriminator | Stage | Layer Type | In→Out | Kernel | Stride | Padding | Params |
| MRAD (per res) | Conv 1 | Conv2d WeightNorm | 1→64 | 7×5 | 2×2 | 3×2 | 2 368 |
| Conv 2 | Conv2d WeightNorm | 64→64 | 5×3 | 2×1 | 2×1 | 61 568 | |
| Conv 3 | Conv2d WeightNorm | 64→64 | 5×3 | 2×2 | 2×1 | 61 568 | |
| Conv 4 | Conv2d WeightNorm | 64→64 | 3×3 | 2×1 | 1×1 | 36 992 | |
| Conv 5 | Conv2d WeightNorm | 64→64 | 3×3 | 2×2 | 1×1 | 36 992 | |
| Conv_post | Conv2d WeightNorm | 64→1 | 3×3 | 1×1 | 1×1 | 578 |
| Discriminator | Stage | Layer Type | In→Out | Kernel | Stride | Padding | Params |
| MRPD (per res) | Conv 1 | Conv2d WeightNorm | 1→64 | 7×5 | 2×2 | 3×2 | 2 368 |
| Conv 2 | Conv2d WeightNorm | 64→64 | 5×3 | 2×1 | 2×1 | 61 568 | |
| Conv 3 | Conv2d WeightNorm | 64→64 | 5×3 | 2×2 | 2×1 | 61 568 | |
| Conv 4 | Conv2d WeightNorm | 64→64 | 3×3 | 2×1 | 1×1 | 36 992 | |
| Conv 5 | Conv2d WeightNorm | 64→64 | 3×3 | 2×2 | 1×1 | 36 992 | |
| Conv_post | Conv2d WeightNorm | 64→1 | 3×3 | 1×1 | 1×1 | 578 |
| Discriminator | Total Params |
| CVaRD | 147 657 |
| CCD | 7 629 |
| MCUD | 7 374 |
| EDD | 15 820 |
| DROD | 15 691 |
| MRAD | 600 198 |
| Stage / Component | Layer Type | In Out | Kernel | Stride | Padding | Heads | Params |
| Pre-processing (NB magnitude/phase feature lift) | |||||||
| Pre-mag convolution | Conv1d | 513 512 | 7 | 1 | 3 | – | 1 839 104 |
| Pre-pha convolution | Conv1d | 513 512 | 7 | 1 | 3 | – | 1 839 104 |
| Pre-mag LayerNorm | LayerNorm | 512 512 | – | – | – | – | 1 024 |
| Pre-pha LayerNorm | LayerNorm | 512 512 | – | – | – | – | 1 024 |
| Swin1DBlock (per block breakdown; dim=512, heads=8, window=8, mlp_ratio=4) | |||||||
| Hyperparameter | Value | Use Case & Rationale |
| Max epochs | 50 | Upper bound on training iterations to ensure convergence while controlling compute budget. |
| Number of GPUs | auto | Hardware-aware scaling: training automatically adapts to available accelerators and enables multi-GPU training when present. |
| Batch size (per GPU) | auto | Per-device minibatch is adjusted to maintain stable memory footprint and consistent throughput across different GPU counts. |
| Random seed | 1234 | Ensures deterministic initialization and repeatable experimental results for fair ablations. |
| Distributed backend | nccl | High-performance multi-GPU communication backend optimized for NVIDIA devices. |
| Initialization URL | tcp://127.0.0.1:54321 | Single-node rendezvous endpoint for distributed process coordination. |
| Hyperparameter | Value | Use Case & Rationale |
| Optimizer (G / D) | AdamW / AdamW | Decoupled weight decay improves optimization stability and generalization under adversarial training dynamics. |
| Generator learning rate | Conservative step size to promote stable convergence of the generator under multi-loss supervision. | |
| Discriminator learning rate | Slightly higher step size to keep discriminators responsive and maintain informative gradients for the generator. | |
| Adam | (0.8, 0.99) | Momentum and second-moment settings tuned for GAN non-stationarity, balancing fast adaptation and variance control. |
| Learning-rate schedule | Exponential | Smooth annealing reduces oscillations late in training and encourages fine-grained refinement near convergence. |
| Learning-rate decay factor | 0.999 | Gentle per-epoch decay to avoid premature stagnation while still improving late-stage stability. |
| Hyperparameter | Value | Use Case & Rationale |
| Mixed precision | enabled (FP16) | Improves throughput and reduces memory; numerically sensitive spectral transforms are kept in full precision for stability. |
| TF32 | enabled | Accelerates matrix operations on compatible GPUs with negligible impact on model quality in practice. |
| cuDNN benchmarking | enabled | Selects efficient kernels to maximize training speed for fixed input shapes. |
| Model compilation | disabled | Prioritizes reproducibility and broad compatibility across environments over potential speed gains. |
| Fused optimizers | disabled | Keeps the training stack portable and consistent across different CUDA/toolchain versions. |
| Hyperparameter | Value | Use Case & Rationale |
| Magnitude reconstruction weight | 45 | Emphasizes accurate log-magnitude recovery to preserve spectral envelope and formant structure. |
| Phase reconstruction weight | 100 | Prioritizes phase consistency (including instantaneous and derivative terms) to reduce temporal smearing and artifacts. |
| Complex-spectrum reconstruction weight | 90 | Encourages coherent complex STFT prediction, improving perceptual fidelity beyond magnitude-only supervision. |
| STFT-consistency weight | 90 | Enforces analysis–synthesis consistency to suppress hallucinated components and stabilize waveform reconstruction. |
| Hyperparameter | Value | Use Case & Rationale |
| Segment size (samples) | 8000 | Fixed-length training segments provide a consistent receptive field and enable efficient batching. |
| High-rate sampling rate (Hz) | 16000 | Target wideband sampling rate defining the output bandwidth for reconstruction. |
| Low-rate sampling rate (Hz) | 4000 | Input narrowband sampling rate defining the degraded bandwidth used as conditioning. |
| Subsampling ratio (HR/LR) | 4 | Specifies the bandwidth expansion factor from narrowband input to wideband target. |
| STFT FFT size | 1024 | Trade-off between frequency resolution and computational efficiency for time–frequency modeling. |
| STFT hop size | 80 | Dense frame sampling improves temporal tracking while maintaining manageable compute. |
| Hyperparameter | Value | Use Case & Rationale |
| Feature dimension | 512 | Hidden width controlling representational capacity for joint magnitude/phase modeling. |
| Pre-projection kernel size | 7 | Enlarged temporal context for initial feature extraction before attention-based processing. |
| Pre-projection input channels | 513 | One-sided STFT bin count corresponding to the chosen FFT size. |
| Layer normalization epsilon | Stabilizes normalization in low-variance regimes and improves training robustness. | |
| Number of lattice coupling blocks | 2 | Two-stage cross-stream coupling to progressively refine magnitude and phase representations. |
| Coupling scalars | (init 1.0) | Learnable interaction strengths enabling adaptive information flow between magnitude and phase streams. |
| Hyperparameter | Value | Use Case & Rationale |
| Swin-1D branch (per lattice block) | ||
| Swin stack depth | 2 blocks | Alternating non-shifted and shifted window attention improves cross-window context while preserving locality. |
| Heads / window / shift | 8 / 8 / 4 | Multi-head attention over short temporal windows; shift enables boundary-crossing interactions with minimal cost. |
| MLP ratio | 4.0 | Expands channel capacity in the feed-forward sublayer for improved nonlinear modeling. |
| Dropout / attention dropout | 0.1 / 0.0 | Regularization via projection/MLP dropout while keeping attention deterministic for stability. |
| LayerScale initialization | Residual scaling to stabilize deep residual learning and mitigate early training instability. | |
| Hyperparameter | Value | Use Case & Rationale |
| Period set | 2, 3, 5, 7, 11 | Captures periodic artifacts across multiple temporal granularities common in neural vocoders and bandwidth expansion. |
| Kernel / stride | 5 / 3 | Controls local pattern sensitivity and progressive temporal downsampling. |
| Channels | 32 128 512 1024 1024 | Increasing capacity enables hierarchical discrimination from local textures to global structure. |
| LeakyReLU slope | 0.1 | Mild negative slope improves gradient flow and avoids dead activations. |
| Normalization | weight norm | Stabilizes discriminator training by constraining effective weight magnitudes. |
| Hyperparameter | Value | Use Case & Rationale |
| MRAD resolutions | (512,128,512) (1024,256,1024) (2048,512,2048) | Multi-scale amplitude evaluation to detect artifacts at short, mid, and long analysis windows. |
| MRPD resolutions | (512,128,512) (1024,256,1024) (2048,512,2048) | Multi-scale phase evaluation to penalize phase inconsistency across resolutions. |
| Base channels | 64 | Provides sufficient capacity for time–frequency feature extraction without excessive compute. |
| STFT window | rectangular | Uses a consistent reference spectrogram definition for multi-resolution comparisons. |
| 2D kernels | (7,5), (5,3), (5,3), (3,3), (3,3) | Controls receptive fields over time–frequency neighborhoods for artifact detection. |
| 2D strides | (2,2), (2,1), (2,2), (2,1), (2,2) | Balances downsampling across frequency and time to maintain discriminative detail. |
| Hyperparameter | Value | Use Case & Rationale |
| CVaRD | ||
| Base channels | 32 | Lightweight separable-convolution backbone to detect waveform artifacts with low overhead. |
| CVaR tail fraction / mode | 0.2 / abs | Emphasizes extreme deviations (two-sided) to improve sensitivity to rare but perceptually salient artifacts. |
| LeakyReLU slope | 0.2 | Stronger negative slope improves gradient propagation in compact discriminators. |
| CCD + Primal–Dual Control | ||
| HF analysis STFT | (1024,80,320) | Matches the main analysis resolution to ensure consistent spectral measurements during constraint enforcement. |
| Hyperparameter | Value | Use Case & Rationale |
| MCUD | ||
| Criteria ( ) | 3 | Aggregates complementary spectral descriptors to approximate perceptual utility beyond adversarial realism alone. |
| Rolloff threshold | 0.85 | Captures spectral energy distribution by measuring the frequency containing 85% of cumulative energy. |
| Utility weights | softmax | Learnable convex combination enables data-driven prioritization among competing criteria. |
| Sample rate | 16000 | Defines Nyquist normalization for frequency-domain criteria such as centroid/rolloff. |
| EDD | ||
| Hyperparameter | Value | Use Case & Rationale |
| DROD | ||
| Base channels | 32 | Efficient per-segment scoring with separable convolutions for robust aggregation. |
| Entropic risk parameter | 1.0 | Tunes robustness: higher values emphasize worst-case segments, improving resilience to hard examples. |
| Aggregation | entropic risk (log-sum-exp) | Smooth worst-case pooling that prioritizes difficult regions without introducing non-differentiable maxima. |
| Loss reweighting | ||
| Evidential / DRO weights (D) | 1.0 / 1.0 | Balances auxiliary discriminators against spectral and waveform discriminators for stable multi-critic training. |
| Module | Hyperparameter | Low / reference / high | Anticipated sensitivity |
| CVaRD | Tail fraction | A smaller value concentrates pooling on a more extreme temporal tail, increasing sensitivity to localized artifacts but also increasing gradient variance. A larger value approaches average pooling and may improve stability at the cost of weaker tail-risk awareness. | |
| CVaRD | Tail-selection mode | right, , left | Absolute selection responds to large activations of either sign and is expected to be the least dependent on feature-map polarity. Right- and left-tail modes may reveal whether the learned activation sign carries consistent risk information, but they may be less robust across layers and random seeds. |
| CCD | High-frequency-band fraction | A smaller fraction restricts the constraint to the highest-frequency bins. A larger fraction also constrains the upper mid-band and is therefore expected to suppress a broader range of spectral amplification, potentially reducing both artifacts and useful high-frequency reconstruction. | |
| CCD | Barrier sharpness | A small produces a smooth penalty and gradual gradients. Increasing makes the transition around the threshold more selective, but an excessively sharp barrier may create abrupt gradients and greater seed-to-seed variability. | |
| CCD | Activation threshold | A lower threshold activates the high-frequency constraint more often and is expected to reduce over-amplification, although it may over-regularize legitimate reconstructed harmonics. A higher threshold is more permissive and may improve spectral detail while allowing more high-frequency artifacts. | |
| CCD | Target mean barrier | A lower target defines a stricter feasible operating point and should increase constraint pressure. A higher target tolerates more high-frequency excess and may improve reconstruction freedom at the expense of weaker artifact control. |
| Phase | ID | Narrowband channel operation | AMR rate | Passband (Hz) | Packet loss | Concealment |
| Training regimes (50 epochs; seeds 1234, 2345, and 3456) | ||||||
| Train | R0 | sinc resampling | – | – | – | – |
| Train | R1 | uniform{resampling, PCM-8, PCM-24} | – | – | – | – |
| Train | R2 | uniform{resampling, G.711 -law, G.711 A-law} | – | – | – | – |
| Train | R3 | uniform{resampling, PCM-8, G.711 -law, G.711 A-law, AMR-NB} | 12.2 kb/s | – | – | – |
| Train | R4 | R3 mixture with telephone filtering and packet loss | 12.2 kb/s | 300–3400 | 3% (20 ms) | repeat previous |