WIPSNet: Deep Learning for Paediatric Wheeze Detection from Overnight Impedance Pneumography
Authors: Felix Oury, Harley Day, Karina Mayoral, Ville-Pekka Seppä, Sejal Saglani, Reiko J. Tanaka
Organizations: Department of Computing, Imperial College London, London, United Kingdom · Department of Bioengineering, Imperial College London, London, United Kingdom · National Heart and Lung Institute and Centre for Paediatrics and Child Health, Imperial College London, London, United Kingdom; Department of Respiratory Paediatrics, Royal Brompton Hospital, London, United Kingdom · Icare Finland Ltd, Vantaa, Finland
Overnight impedance pneumography (IP) is used to monitor paediatric respiratory health. Its current clinical readout, the Expiratory Variability Index (EVI), compresses each IP recording into a single scalar and achieves an AUC of 0.633 for night-level wheeze classification. We introduce Wheeze Impedance Pneumography Scalogram Network (WIPSNet), a 3D ResNet operating on stacked continuous wavelet transform scalograms of overnight IP signals. On a 15-patient cohort (60 nights, 281 hours), WIPSNet achieves an AUC of 0.783±0.026, outperforming EVI, a state-space model (Mamba), and two modern sleep-staging architectures. Performance peaks at a volumetric depth corresponding to 32 minutes of temporal context, suggesting that multi-scale temporal aggregation is important for modelling nocturnal respiratory dynamics. Overall, these results indicate that structured time-frequency representations combined with 3D convolutional architectures provide an effective approach for learning from long, irregular physiological time series.
Figures & tables
Figure 1: WIPSNet framework and architecture. (a) The IP signal is filtered, transformed into scalograms, grouped into 3D volumes, and passed through a 3D CNN; volume-level probabilities are aggregated into a night-level prediction. (b–d) Raw signal, processed signal, and 4-minute scalogram. (e) Volumetric input construction: D consecutive scalograms stacked along the depth axis. (f) The bottleneck 3D ResNet backbone.
Model
Input & backbone
AUC
AUPRC
WIPSNet
CWT + 3D CNN
.783±.026
.634±.047
AttnSleep
Raw + MRCNN-attn
.764±.030
.603±.048
catch22+XGBoost
22 canonical TS features
.704±.002
.481±.002
Mamba
Raw + SSM
.689±.047
.509±.060
EVI †
Clinical scalar
.633
.433
SleePyCo
Raw + Pyramid-Transformer
.577±.152
.444±.101
Table 1 : Night-level wheeze classification results. Mean ± std over 5 seeds, 15-fold patient-level LOOCV. AUPRC random baseline at our class prevalence is 0.35. Mamba and Mamba-Scalogram both use a 32-min window. † Deterministic; 57/60 valid nights.
Figure 2 : WIPSNet’s depth-scaling behaviour. Both AUC and AUPRC peak at D=8 (32 min of temporal context) and decline at longer windows, indicating an optimal macro-timescale rather than monotone improvement. Error bars: standard deviation over 5 seeds. Top axis: temporal context in minutes.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Search Space
Selected Value
Scalogram window length
2–16 min
60,000 (4 min)
Volume depth D
1–10 scalograms
8
Stride (scalogram)
1– D scalograms
8 ( =D , non-overlapping)
Backbone
ResNet / EfficientNet
ResNet
Batch size
{8, 16, 32, 64}
32
Learning rate
10−5 – 10−2
3.00×10−4
Appendix
Table 2 : WIPSNet hyperparameter search space and selected values.
Parameter
Search Space
Selected Value
Context window W (4-min segments)
{1, 2, 4, 8} sweep
8 ( =32 min)
Stride (segments)
=W (non-overlapping)
8
Per-segment length
—
5,000 samples (4 min) 3 3 3 The raw 250 Hz IP signal is downsampled 12× to ≈ 20.8 Hz before Mamba, so a 4-min segment is 5,000 samples and a 32-min window is 8×5,000=40,000 samples. This keeps the SSM sequence length tractable; on the full-rate 480,000-sample window Mamba collapses to near-chance.
Sequence length
W× seg. length
40,000 samples (32 min)
dmodel
16–192
32
dstate
8–48
32
Appendix
Table 3 : Mamba SSM hyperparameter search space and selected values. The context window W (number of consecutive 4-min segments) was swept separately, analogous to the WIPSNet volume depth D , while the architecture was tuned by Optuna.
Model
Input & backbone
AUC
AUPRC
WIPSNet-CWT ( D=8 )
CWT + 3D CNN
.783±.026
.634±.047
AttnSleep
Raw 1D + MRCNN + attention
.764±.030
.603±.048
WIPSNet-CWT ( D=4 )
CWT + 3D CNN
.747±.032
.556±.035
WIPSNet-CWT ( D=1 )
CWT + 3D CNN
.744±.022
.527±.026
WIPSNet-CWT ( D=32 )
CWT + 3D CNN
.744±.042
.546±.049
WIPSNet-CWT ( D=2 )
CWT + 3D CNN
.742±.035
.558±.044
Appendix
Table 4 : All models and ablations, sorted by AUC. Mean ± std over 5 seeds for DL models. AUC is the primary ranking metric; AUPRC reflects performance on the minority (wheeze) class under class imbalance. Depth-sweep rows ( D=8 ) and front-end rows share WIPSNet’s backbone; the headline model is WIPSNet-CWT ( D=8 ).
Department of Pediatrics, Baylor College of Medicine, Houston, TX, USA · Jan and Dan Duncan Neurological Research Institute, Texas Children’s Hospital, Houston, TX, USA · Institute for Analysis and Numerics, University of Münster, Germany +1
Department of Medical Engineering and Technomathematics, FH Aachen University of Applied Sciences, 52428 Jülich, Germany · Department of Information and Computing Sciences, Utrecht University, Utrecht, The Netherlands · Institute for Data-Driven Technologies, FH Aachen University of Applied Sciences, 52428 Jülich, Germany