This paper presents and evaluates a Deep Learning-based (DL-based) Signal Quality Assessment (SQA) model to distinguish between clean and noisy ambulatory Electrocardiograms (ECG). The model is trained on Copenhagen Center for Health Technology-Contextualized Arrhythmia Database (CACHET-CADB), which, to the best of our knowledge, is the first ambulatory ECG database with both physical and patient-reported contextual data. The model shows stable performance on different databases such as MIT-databases and the latest PyhsioNet/Cinc Challenge 2021 databases. Subsequently, the paper demonstrates how complicated ECG noise can be investigated by the SQA model and the physical contextual data.
Figures & tables
Category
NSR
AFIB
Other
Noise
Number of signals
615
747
19
221
Considered class
Clean
Clean
Clean
Noisy
TABLE I: ECG rhythms from CACHET-CADB.
Fig. 1: Scalogram of a 10 s NSR
Fig. 2: Scalogram of a 10 s noise
Fig. 3: Architecture of the final model. The three input images are identical. The red neurons in Dense layer 1 are prohibited by a dropout probability.
Hyper-parameter
Value
Dense layer 1 size
20
Learning rate
1⋅10−5
Dropout probability
0.6
Batch size
20
Epochs
40
TABLE II: Optimized hyper-parameters of the DL model.
Attribute
Parameter
0: Unknown
1: Lying supine
2: Lying left
3: Lying prone
Body position
4: Lying right
5: Upright
TABLE III: The contextualized parameters presented in [ 1 ] .
Actual class
1,602 labelled signals in total.
True
False
True
TP = 215
FP = 128
Predicted positives = 343
Prediction
False
FN = 6
TN = 1,253
Predicted negatives = 1,259
Se = 97.3 %
Sp = 90.7 %
Acc = 91.6%
TABLE IV: Confusion matrix of the model performance on CACHET-CADB test set.
Database
Record amount
Duration (h)
Percentage of predicted clean signals (%)
MIT-BIH-NSR
3
48
98.2
MIT-BIH-Arrhythmia
48
24
76.9
MIT-BIH-Noise Stress Test
3
1.5
2.2
TABLE V: Prediction on different MIT-BIH databases.
Category
Number of signals
Number of predicted clean signals
Percentage of predicted clean signals (%)
Sinus Bradycardia (SB)
3,889
3,605
92.7
Normal Sinus Rhythm (NSR/SR)
1,826
1,661
91.0
Atrial Fibrillation (AFIB)
1,780
1,171
65.8
Sinus Tachycardia (ST)
1,568
1,305
83.2
Supraventricular Tachycardia (SVT)
587
345
58.8
Atrial Flutter (AF)
445
280
62.9
TABLE VI: DL model performance on the Chapman-Shaoxing database. The worst performances are marked with orange.
Category*
Number of signals
Number of predicted clean signals
Percentage of predicted clean signals (%)
Sinus Bradycardia (SB)
11,919
11,059
92.8
Normal Sinus Rhythm (NSR/SR)
6,058
5,489
90.6
Atrial Flutter (AF)
5,164
3,498
67.7
Sinus Tachycardia (ST)
3,820
3,182
83.3
Sinus arrhythmia
1,294
1,180
91.2
Premature Atrial Contraction (PAC)
975
681
69.8
TABLE VII: DL model performance on the Ningbo arrhythmia database. The worst performances are marked with orange.
Electrocardiogram (ECG) recordings are corrupted by non-stationary noise sources that degrade diagnostic reliability, particularly in ambulatory and long-duration recordings. Deep learning denoisers exist, but convolutional architectures are limited by their receptive field, transformer-based models scale quadratically with sequence length, and diffusion-based approaches incur prohibitive inference cost. We propose a Mamba-augmented model that inserts selective state-space blocks at the convolutional bottleneck, combining local feature extraction with long-range temporal modeling at linear complexity. We comprehensively evaluate the proposed model with respect to reconstruction fidelity, noise robustness, recording-length scaling, and downstream diagnostic classification across over 40 pathology classes. On synthetic and real datasets, our model achieves the highest SNR and lowest RMSE, with the Mamba advantage increasing with sequence length and in low-SNR regimes. On classification with two independent classifiers, the proposed Mamba-based models achieve the best macro AUROC among all denoisers and improve over their convolutional base models. Calibration is more nuanced and classifier-dependent: denoising improves Binary Cross-Entropy and Brier score on Inception1D but often fails to beat the noisy input on ResNet1D-Wang, and the lead-specific Mamba variant is the only denoiser to improve both calibration metrics over the noisy baseline on both classifiers. Per-class analysis reveals a morphology-dependent benefit: Mamba substantially improves ST/T-change diagnoses, which depend on broad, context-sensitive waveforms.
Basile Morel, Samuel Ruiperez-Campillo, Andreas P. Streich +2
Department of Computer Science at ETH Zurich, Universitätstrasse 6, 8092 Zürich, Switzerland
ECG-Language Models (ELMs) extend recent advances in Multimodal Large Language Models (MLLMs) to automated ECG interpretation. However, most existing ELMs inherit Vision-Language Model (VLM) design choices and rely on pretrained ECG encoders, introducing substantial architectural and training complexity. Inspired by encoder-free VLMs, we introduce ELF, a family of three encoder-free ELMs that remain competitive with, and often outperform, prior state-of-the-art ELMs across two datasets despite substantially simpler architectures and training pipelines. All code and data are available at github.com/ELM-Research/ECG-Language-Models.
William Han, Tony Chen, Chaojing Duan +6
Carnegie Mellon University, USA · Allegheny Health Network, USA · University of California, Los Angeles, USA +1
Self-supervised electrocardiogram (ECG) models are often trained on a few seconds of ECG signal and, increasingly, on discretized token sequences. It remains unclear whether these choices sacrifice information needed for rhythm inference and longitudinal consistency in real-world ambulatory recordings. We present a controlled study on the Icentia11k single-lead dataset that varies (i) the input horizon (16 seconds, 1 minute, 5 minutes, and 10 minutes) and (ii) the front-end representation (continuous convolutional patch embeddings vs. fixed vector-quantized tokens), while holding the Transformer backbone and training protocol constant. Representations are assessed by downstream abnormal rhythm detection and by patient-level retrieval that probes cross-session stability. Our results show that increasing temporal context beyond 16-second snapshots yields stronger transfer and higher retrieval accuracy, with the strongest performance achieved by the 5- and 10-minute models, indicating improved capture of slow-varying rhythm dynamics and individual-specific structure. Across all evaluated horizons, continuous patch embeddings outperform discretized tokens, suggesting that quantization can discard clinically relevant waveform detail. These findings motivate ECG foundation models that emphasize extended context and continuous encoders for clinical prediction and similarity-based applications. Our code and pretrained models are publicly available at https://github.com/muha-0/ecg-ssl-representation-learning.
Ahmed Sameh, Ramzi Al-Sharawi, Yogatheesan Varatharajah
Computer Science & Engineering, University of Minnesota Twin Cities, Minneapolis, MN 55455. · Robotics, University of Minnesota Twin Cities, Minneapolis, MN 55455.