Label-Efficient Deep Learning for ECG Delineation: A Multi-Dataset Benchmark against Widely Used Delineation Tools
Authors: Jeonghwa Lim, Minje Park, Yeongyeon Na, Yujin Eom, Soyeon Lim, Young Ho Lee, Yu Jeong Kim, Sunghoon Joo, +1 more
Organizations: VUNO Inc., Seoul, South Korea · Department of AI Mobility Engineering, Ajou University, Suwon, South Korea · C&Thoth Co., Ltd., Seoul, South Korea · Department of Biomedical Sciences, Chonnam National University Graduate School, and the Department of Cardiovascular Medicine, Chonnam National University Hospital, Gwangju, South Korea · Department of Cardiovascular Medicine, Chonnam National University Hospital, Gwangju, South Korea · Department of Cardiovascular Medicine, Chonnam National University Hospital, and the Department of Internal Medicine, Chonnam National University Medical School, Gwangju, South Korea
Electrocardiogram (ECG) delineation, the identification of waveform boundaries, is a foundational step that translates raw ECG signals into clinically interpretable measurements. Deep learning has advanced this task but remains dependent on costly expert annotations. Label-efficient strategies such as self-supervised pretraining and semi-supervised learning are expected to ease this burden, yet it remains unclear whether they yield reliable delineation and whether the deep models they produce outperform the delineation tools used in practice. We address this in two stages. First, comparing self-supervised objectives with supervised or semi-supervised fine-tuning across one internal and four external datasets, we find that pretraining helps but the objective matters, and that the value of semi-supervised fine-tuning depends on the pretraining objective. Second, we benchmark the selected deep learning model against widely used open-source (NeuroKit2, Prominence, ECGdeli) and commercial (CalECG) tools using three complementary metrics. The model ranks best on every metric and dataset, outperforming the strongest tool by a clear margin on the rhythm-diverse set (mIoU 71.3 vs. 54.8%; averaged point-wise sensitivity 92.6 vs. 76.4%), and degrades the least from sinus to arrhythmia. A rhythm-stratified and point-wise analysis further characterizes the distinctive behavior of each tool, yielding practical guidance for tool selection. These results provide systematic, multi-dataset evidence that self-supervised pretraining is effective for ECG delineation and enables a label-efficiently trained deep learning model to outperform widely used delineation tools by leveraging abundant unlabeled data. This supports adopting such models in diverse, real-world clinical settings.
Figures & tables
Fig. 1: Study overview. Stage 1 pretrains an encoder on unlabeled ECG data with self-supervised learning (e.g., Masked Autoencoder) and fine-tunes it with a segmentation decoder (e.g., FCN) under supervised and semi-supervised learning (e.g., Mean Teacher). Stage 2 benchmarks the resulting model against existing delineation tools on multiple datasets using three complementary metrics: mIoU, measurement error on the PR, QRS, and QT intervals, and fiducial point localization accuracy.
Role
Source
#Subjects
#ECGs
Duration (labeled)
Sample rate
Lead type
#Samples
Training
QTDB
105
105
5.9–253.6 s
250 Hz
2-lead
718
ISP
499
499
10 s
1000 Hz
12-lead
5,988
PTB-XL
18,885
21,837
— (unlabeled)
500 Hz
12-lead
262,044
Evaluation
LUDB
200
200
10 s
500 Hz
12-lead
2,369
Zhejiang
334
334
1.3–7.1 s
2000 Hz
12-lead
4,008
RDB
2,399
2,399
10 s
500 Hz
12-lead
28,788
TABLE I: Characteristics of the ECG datasets used for developing and evaluating the delineation model. Lead configurations range from 12-lead (six limb and six precordial leads) to 6-lead (limb only) and 2-lead (arbitrary pairs, e.g., MLII and V1).
mIoU (%) ↑
Summary
Pretraining
Fine-tuning
Internal
LUDB
Zhejiang
RDB
Avg. (%) ↑
Rank ↓
–
Supervised
75.2±0.2
64.8±0.8
53.3±2.4
60.5±0.5
63.5
3.50
Semi-supervised
81.6±0.2
70.0±5.9
66.0±5.3
67.9±0.9
71.4
MAE ∗
Supervised
82.8±0.3
71.3±2.3
75.2±1.5
71.9±1.2
75.3
1.25
Semi-supervised
82.3±0.5
72.8±5.3
74.8±2.1
71.1±0.4
75.2
MoCo ∗
Supervised
82.2±0.1
69.5±3.3
66.2±3.7
71.6±1.5
72.4
1.75
TABLE II: Benchmarking results of learning strategies for ECG delineation. Per-dataset entries are mIoU, reported as mean ± standard deviation over three seeds, where Internal denotes the QTDB + ISP test split. Under Summary, Avg. is the mIoU averaged across the four datasets for each fine-tuning strategy, and Rank is the average rank among the four pretraining strategies, aggregated over fine-tuning and datasets. The best value in each column is shown in bold .
mIoU (%) ↑
Averaged interval error (ms) ↓
Averaged point-wise sensitivity (%) ↑
Method
Internal
LUDB
Zhejiang
RDB
Internal
LUDB
Zhejiang
RDB
Internal
LUDB
Zhejiang
RDB
Avg. Rank ↓
NeuroKit2 (CWT)
41.3
44.9
7.3
40.0
70.9
108.3
72.8
30.2
59.4
59.9
15.7
54.9
5.8
NeuroKit2 (DWT)
43.4
45.6
8.6
38.6
43.6
45.6
52.7
40.5
60.3
62.2
17.8
58.5
5.2
ECGdeli
67.5
59.4
41.9
54.8
24.6
29.3
30.4
30.0
82.4
80.6
56.3
76.4
3.4
Prominence
60.4
63.9
42.7
51.7
28.1
22.3
35.9
23.0
78.1
82.9
54.6
73.7
3.6
DL delineator (sup)
83.0
71.6
73.5
71.1
11.5
19.0
16.3
22.2
96.1
86.6
82.9
92.6
1.8
TABLE III: Comparison of the DL delineator with open-source delineation tools. The best value in each column is in bold , and the second-best is underlined . Avg. Rank is the mean rank across all columns. CalECG , configured for interval outputs only, is compared separately on interval error in the Supplementary Material (Table S10 ).
Method
Average
PR
QRS
QT
NeuroKit2 (CWT)
97.3
207.4
32.1
52.3
NeuroKit2 (DWT)
51.3
45.3
17.0
91.5
ECGdeli
32.9
47.0
18.0
33.6
CalECG
37.9
41.1
32.7
40.0
Prominence
25.4
29.6
25.0
21.7
DL delineator (sup)
15.0
13.8
12.2
19.1
TABLE IV: Comparison of the DL delineator on mECGDB with open-source and commercial delineation tools, evaluated by mean absolute error (ms) for the PR interval, QRS duration, QT interval, and their average. The best value is in bold and the second-best is underlined .
Dataset
Method
P on
P off
QRS on
QRS off
T on
T off
Total
LUDB
NeuroKit2 (CWT)
43.7
43.1
80.8
72.2
56.4
63.4
59.9
NeuroKit2 (DWT)
73.3
63.8
66.2
81.8
32.5
55.8
62.2
ECGdeli
71.8
87.2
97.3
85.1
70.7
71.7
80.6
CalECG
74.0
-
67.7
69.1
-
61.5
68.1
Prominence
79.2
80.8
91.9
85.8
79.1
80.7
82.9
DL delineator (sup)
80.4
81.3
99.1
98.9
77.8
81.6
86.6
TABLE V: Comparison of the DL delineator on external test sets ( LUDB and Zhejiang ) with open-source and commercial delineation tools, evaluated by point-wise sensitivity (%) per fiducial point. ”Total” is the average over the available onset/offset points. The best value is in bold and the second-best is underlined .
Subgroup performance
Metric
Method
Sinus
Arrhythmia
Degradation
mIoU (%) ↑
ECGdeli
66.3
38.7
41.6%
Prominence
64.4
33.2
48.5%
DL delineator (sup)
76.1
64.9
14.7%
DL delineator (semi)
76.6
64.2
16.2%
Averaged interval error (ms) ↓
ECGdeli
28.9
40.1
38.8%
TABLE VI: Rhythm-stratified subgroup performance on RDB for the sinus and arrhythmia groups, with the relative performance degradation from sinus to arrhythmia (smaller is better). The best value is in bold and the second-best is underlined .
Fig. 2: Qualitative examples of ECG delineation results across methods and datasets. The rhythm classes shown reflect the characteristics of each dataset. Rows are grouped into the reference annotation (GT), rule-based tools, and the DL delineator ; shaded bands mark each method’s detected P (blue), QRS (orange), and T (green) waves on the same input signal. SR, sinus rhythm; PVC, premature ventricular contraction; AFIB, atrial fibrillation; SVT, supraventricular tachycardia.
Source
Train
Validation
Test
Total
QTDB
422
148
148
718
ISP
3,792
1,272
924
5,988
Internal ( QTDB + ISP )
4,214
1,420
1,072
6,706
TABLE S1: The detail number of the Internal ( QTDB + ISP ) set sample.
Dataset
Rhythm
#ECGs
LUDB
Sinus rhythm
143
Sinus tachycardia
4
Sinus bradycardia
25
Sinus arrhythmia
8
Irregular sinus rhythm
2
Abnormal rhythm (AFib, flutter)
18
TABLE S2: Rhythm composition of the LUDB and Zhejiang used for external test.
Group
Rhythm
#ECGs
Sinus
Sinus rhythm
400
Sinus bradycardia
400
Sinus tachycardia
140
Sinus irregularity
399
Arrhythmia
Atrial flutter
400
Atrial fibrillation
400
TABLE S3: Rhythm composition of the RDB used for the rhythm-stratified subgroup analysis. Recordings are grouped into sinus and arrhythmia categories.
Component
Setting
Framework
Mean Teacher (MT)
Teacher EMA decay
0.99
Consistency loss weight
1.0
Weak augmentation
Random resized cropping
Strong augmentation
RandAugment( N=3,p=0.5 )
Candidate transforms
Powerline noise
TABLE S4: Semi-supervised learning configuration.
Hyperparameter
MAE
MoCo
MERL
Epochs
800
300
50
Warm-up epochs
40
40
5
Batch size
2048
2048
2048
Learning rate
1.2e-3
1.2e-3
2.0e-4
Weight decay
0.05
0.1
1e-5
(β1,β2)
(0.9, 0.95)
(0.9, 0.999)
(0.9, 0.999)
TABLE S5: Optimization hyperparameters for self-supervised pretraining: MAE , MoCo , and MERL . All algorithms are trained with the AdamW optimizer and cosine learning rate scheduling.
Hyperparameter
Value
MAE
Masking ratio
0.75
Decoder embedding dimension
96
Decoder depth
4
Decoder attention heads
3
MoCo
TABLE S6: Model architecture and algorithm-specific hyperparameters for self-supervised pretraining.
Method
Time (ms) ↓
NeuroKit2 (CWT)
56.3 ± 11.9
NeuroKit2 (DWT)
63.6 ± 11.1
ECGdeli
141.0 ± 57.5
Prominence
20.0 ± 6.0
DL delineator
17.1 ± 1.9
TABLE S7: Computation time analysis on LUDB , reported as a mean ± standard deviation. The best value is bolded.
Pretraining
Fine-tuning
Internal
LUDB
Zhejiang
mECGDB
RDB
–
Supervised
16.2±0.3
21.1±0.7
33.1±2.0
26.1±0.9
23.4±0.3
Semi-supervised
12.8±0.4
20.8±2.4
21.7±3.0
16.2±0.6
29.8±4.3
MAE
Supervised
11.6±0.2
19.2±1.7
16.5±1.2
15.0±0.2
20.9±1.2
Semi-supervised
11.7±0.3
18.4±2.0
15.9±0.6
15.7±0.7
22.7±1.0
MoCo
Supervised
12.0±0.3
21.3±2.5
21.3±1.0
15.2±0.8
22.4±1.2
Semi-supervised
11.8±0.3
17.2±0.8
16.9±0.8
15.4±0.7
22.3±0.5
TABLE S8: Benchmarking results of learning strategies for ECG delineation. Per-dataset entries are averaged interval error (ms), reported as mean ± std over three seeds, where Internal denotes the QTDB + ISP test split. The best value in each column is shown in bold .
Pretraining
Fine-tuning
Internal
LUDB
Zhejiang
RDB
–
Supervised
91.3±0.2
82.5±1.2
66.5±2.6
87.1±0.7
Semi-supervised
95.4±0.2
86.5±7.0
76.6±6.2
90.8±0.6
MAE
Supervised
96.1±0.2
86.4±3.0
85.4±2.2
93.3±0.6
Semi-supervised
96.2±0.3
89.2±3.3
85.9±1.5
92.6±0.3
MoCo
Supervised
95.7±0.1
85.1±3.5
77.5±4.2
92.4±1.0
Semi-supervised
96.1±0.2
90.3±3.5
84.5±0.7
92.9±1.0
TABLE S9: Benchmarking results of learning strategies for ECG delineation. Per-dataset entries are averaged point-wise sensitivity (%), reported as mean ± std over three seeds, where Internal denotes the QTDB + ISP test split. The best value in each column is shown in bold .
Averaged interval error (ms) ↓
Method
Internal
LUDB
Zhejiang
CalECG
37.9
20.1
58.2
DL delineator (sup)
11.5
19.0
16.3
DL delineator (semi)
12.0
16.8
15.3
TABLE S10: Comparison of the DL delineator with CalECG on averaged interval error. The best value is in bold and the second best is underlined .
Sensitivity (%)
PPV (%)
Dataset
Method
40 ms
150 ms
40 ms
150 ms
Internal
NeuroKit2 (CWT)
59.4
73.7
66.1
82.1
NeuroKit2 (DWT)
60.3
88.4
60.5
88.0
ECGdeli
82.4
97.0
77.3
90.8
Prominence
78.1
95.9
74.2
91.1
DL delineator (sup)
96.1
99.0
95.6
98.6
TABLE S11: Comparison of the DL delineator with open-source delineation tools, evaluated by averaged point-wise total sensitivity and positive predictive value (PPV) (%) at tolerances of 40 and 150 ms. The best value in each column is in bold .
Dataset
Method
P on
P off
QRS on
QRS off
T on
T off
Total
LUDB
NeuroKit2 (CWT)
43.5
42.8
87.4
93.8
67.5
76.0
68.5
NeuroKit2 (DWT)
62.4
53.6
75.1
92.3
32.7
56.5
62.1
ECGdeli
55.1
66.9
96.8
84.6
63.2
64.1
71.8
CalECG
83.8
-
95.0
96.9
-
84.5
90.1
Prominence
64.6
65.9
91.0
84.9
79.0
80.7
77.7
DL delineator (sup)
88.1
89.1
99.2
99.0
81.4
85.4
90.4
TABLE S12: Comparison of the DL delineator on external test sets ( LUDB and Zhejiang ) with open-source and commercial delineation tools, evaluated by point-wise positive predictive value (PPV, %) per fiducial point. ”Total” is the average over the available onset/offset points. The best value is in bold .
Dataset
Method
P on
P off
QRS on
QRS off
T on
T off
LUDB
NeuroKit2 (CWT)
-7.6 ± 9.4
0.1 ± 9.5
-8.8 ± 12.3
-12.0 ± 13.3
5.0 ± 17.4
6.1 ± 10.9
NeuroKit2 (DWT)
3.5 ± 17.6
-3.0 ± 17.3
-3.6 ± 13.7
-3.1 ± 13.4
17.7 ± 20.1
-10.0 ± 13.2
ECGdeli
-13.0 ± 11.3
16.8 ± 11.4
-7.2 ± 13.9
16.2 ± 16.0
-10.4 ± 19.9
4.5 ± 15.9
CalECG
0.3 ± 11.5
-
8.1 ± 13.5
-2.3 ± 11.8
-
5.3 ± 12.1
Prominence
1.9 ± 12.8
3.1 ± 13.9
-0.0 ± 15.9
-11.4 ± 18.1
2.8 ± 18.2
13.2 ± 14.8
DL delineator (sup)
-5.4 ± 14.3
7.6 ± 15.1
-8.4 ± 10.1
10.4 ± 11.4
4.0 ± 20.1
12.5 ± 14.1
TABLE S13: Comparison of the DL delineator on LUDB and Zhejiang external test sets with widely used open-source and commercial delineation tools, evaluated by point-wise localization error (ms) per fiducial point at the 40 ms tolerance. The signed mean error and its standard deviation are reported.
Sinus group
Arrhythmia group
Metric
Method
SR
SB
ST
SI
Avg.
AF
AFIB
AT
SVT
Avg.
Deg.%
mIoU (%) ↑
ECGdeli
67.8
67.2
63.2
66.9
66.3
38.8
42.1
45.3
28.8
38.7
41.5
Prominence
67.2
65.4
58.9
66.3
64.4
33.1
38.1
38.6
22.8
33.2
48.6
DL delineator (sup)
76.4
77.0
75.0
75.9
76.1
61.1
67.6
63.5
67.4
64.9
14.7
DL delineator (semi)
77.3
78.0
74.2
76.9
76.6
59.4
67.8
63.4
66.3
64.2
16.2
Averaged interval error (ms) ↓
ECGdeli
28.3
30.1
28.4
28.6
28.9
34.1
29.8
34.4
62.0
40.1
65.8
TABLE S14: Rhythm-stratified subgroup analysis on the RDB , comparing the DL delineator with the two strongest tools (Prominence, ECGdeli). Rhythms are grouped into sinus (SR, SB, ST, SI) and arrhythmia (AF, AFIB, AT, SVT). For each metric, we report the group mean performance and its relative degradation (Deg.%) from sinus to arrhythmia. The best value is in bold .
Fig. S1: Qualitative P/QRS/T delineation on a representative 2.5 s Lead II segment from LUDB . Rows are grouped into the reference annotation (GT), rule-based tools, and the DL delineator ; shaded bands mark each method’s detected P (blue), QRS (orange), and T (green) waves on the same input signal.
Fig. S2: Qualitative P/QRS/T delineation on a representative 2.5 s Lead II segment from Zhejiang . Rows are grouped into the reference annotation (GT), rule-based tools, and the DL delineator ; shaded bands mark each method’s detected P (blue), QRS (orange), and T (green) waves on the same input signal.
Fig. S3: Qualitative P/QRS/T delineation on a representative 2.5 s Lead II segment from RDB . Rows are grouped into the reference annotation (GT), rule-based tools, and the DL delineator ; shaded bands mark each method’s detected P (blue), QRS (orange), and T (green) waves on the same input signal. SR, sinus rhythm; AFIB, atrial fibrillation; SVT, supraventricular tachycardia.
Electrocardiogram (ECG) delineation identifies the boundaries of P waves, QRS complexes, and T waves, providing structural annotations that can guide AI models in learning to interpret ECGs. However, training accurate delineation models requires manual annotations that are scarce and time-consuming to obtain. Recent work addresses this limitation through semi-supervised learning (SSL), but the design of the architecture, particularly the decoder, has received less attention. To this end, we propose R-U-Net, an ECG delineation model that pairs a ResNet-18 encoder with a U-Net decoder. On SemiSegECG, R-U-Net outperforms the strongest evaluated ResNet-18 + fully convolutional network (FCN) head baseline in each of the 16 in-domain settings by 3.3-13.0 mIoU and achieves 82.6 mIoU in the cross-domain setting, an improvement of 8.1 mIoU. Controlled ablations show that decoder design contributes more to performance gains than the evaluated SSL methods, motivating further exploration of architectures for ECG delineation. All code is open-source at github.com/ELM-Research/ECG-Delineation.
Joseph Scharpf, William Han, Chaojing Duan +3
Carnegie Mellon University · Allegheny Health Network · University of Colorado
While Deep Learning (DL) enhances automated electrocardiogram (ECG) analysis, clinical deployment is hindered by class imbalance and the generalization gap. This paper presents HeartBeatAI, a deep learning framework combining domain generalization, multi-scale feature aggregation, and clinical explainability for robust 12-lead ECG classification. Moving beyond image-based paradigms, HeartBeatAI integrates a Squeeze-and-Excitation (SE) ResNet to isolate diagnostic leads alongside a Multi-Layer Concentration Pipeline to capture macro-rhythm and micro-morphological anomalies. To mitigate domain shift, the framework employs MixStyle regularization and Label Smoothing. Rigorous benchmarking across four large-scale datasets using intra-source and Leave-One-Domain-Out (LODO) protocols demonstrates high performance (98% Macro F1-score) under intra-source conditions. However, LODO evaluations reveal significant degradation in detecting rare anomalies, highlighting a persistent challenge in cross-institutional deployment.
Shubham Gupta, Nikhil Panwar, Partha Pratim Roy
Department of Computer Science and Engineering, Indian Institute of Technology (ISM) Dhanbad, Dhanbad, Jharkhand, India. · Department of Computer Science and Engineering, Indian Institute of Technology Roorkee, Roorkee, Uttarakhand, India.
Electrocardiogram (ECG) arrhythmia classification remains challenging due to signal variability, noise, limited labeled data, and the difficulty in achieving both accuracy and efficiency in models. While self-supervised learning reduces label dependency, most methods target either global contextual features or local morphological patterns, but rarely implement hierarchical multi-scale feature extraction. ECG signals require architectures that simultaneously capture fine-grained beat-level morphology and broader rhythm-level dependencies with computational efficiency. To overcome this limitation, this paper proposes the Electrocardiogram Neighborhood Attention Transformer (ECG-NAT), a novel self-supervised learning approach tailored for multi-lead ECG classification. Our two-stage approach begins with generative pretraining, using a masked autoencoder to reconstruct partially masked ECG signals across multiple diverse datasets, enabling the model to learn robust, domain-invariant representations from unlabeled data. This is followed by discriminative fine-tuning with a dual-loss function that combines supervised contrastive and cross-entropy losses, aligning representation learning with label prediction. The hierarchical attention mechanism efficiently captures multi-scale temporal features from localized beat morphology to broader rhythm patterns at low computational cost. ECG-NAT achieves robust performance on benchmark datasets, with 88.1% accuracy using only 1% labeled data, demonstrating strong efficacy in low-resource settings. The framework combines superior classification performance with computational efficiency, making it practical for real-time ECG diagnosis. The code will be made available upon acceptance at: https://github.com/Mahsagazeran/ECG-NAT.
Department of Computer Engineering, University of Kurdistan, Sanandaj, Iran · Department of Mathematics and Operational Research, University of Mons, Mons, Belgium