Label-Efficient Deep Learning for ECG Delineation: A Multi-Dataset Benchmark against Widely Used Delineation Tools
Authors: Jeonghwa Lim, Minje Park, Yeongyeon Na, Yujin Eom, Soyeon Lim, Young Ho Lee, Yu Jeong Kim, Sunghoon Joo, +1 more
Organizations: VUNO Inc., Seoul, South Korea · Department of AI Mobility Engineering, Ajou University, Suwon, South Korea · C&Thoth Co., Ltd., Seoul, South Korea · Department of Biomedical Sciences, Chonnam National University Graduate School, and the Department of Cardiovascular Medicine, Chonnam National University Hospital, Gwangju, South Korea · Department of Cardiovascular Medicine, Chonnam National University Hospital, Gwangju, South Korea · Department of Cardiovascular Medicine, Chonnam National University Hospital, and the Department of Internal Medicine, Chonnam National University Medical School, Gwangju, South Korea
Electrocardiogram (ECG) delineation, the identification of waveform boundaries, is a foundational step that translates raw ECG signals into clinically interpretable measurements. Deep learning has advanced this task but remains dependent on costly expert annotations. Label-efficient strategies such as self-supervised pretraining and semi-supervised learning are expected to ease this burden, yet it remains unclear whether they yield reliable delineation and whether the deep models they produce outperform the delineation tools used in practice. We address this in two stages. First, comparing self-supervised objectives with supervised or semi-supervised fine-tuning across one internal and four external datasets, we find that pretraining helps but the objective matters, and that the value of semi-supervised fine-tuning depends on the pretraining objective. Second, we benchmark the selected deep learning model against widely used open-source (NeuroKit2, Prominence, ECGdeli) and commercial (CalECG) tools using three complementary metrics. The model ranks best on every metric and dataset, outperforming the strongest tool by a clear margin on the rhythm-diverse set (mIoU 71.3 vs. 54.8%; averaged point-wise sensitivity 92.6 vs. 76.4%), and degrades the least from sinus to arrhythmia. A rhythm-stratified and point-wise analysis further characterizes the distinctive behavior of each tool, yielding practical guidance for tool selection. These results provide systematic, multi-dataset evidence that self-supervised pretraining is effective for ECG delineation and enables a label-efficiently trained deep learning model to outperform widely used delineation tools by leveraging abundant unlabeled data. This supports adopting such models in diverse, real-world clinical settings.
Figures & tables
Fig. 1: Study overview. Stage 1 pretrains an encoder on unlabeled ECG data with self-supervised learning (e.g., Masked Autoencoder) and fine-tunes it with a segmentation decoder (e.g., FCN) under supervised and semi-supervised learning (e.g., Mean Teacher). Stage 2 benchmarks the resulting model against existing delineation tools on multiple datasets using three complementary metrics: mIoU, measurement error on the PR, QRS, and QT intervals, and fiducial point localization accuracy.
Role
Source
#Subjects
#ECGs
Duration (labeled)
Sample rate
Lead type
#Samples
Training
QTDB
105
105
5.9–253.6 s
250 Hz
2-lead
718
ISP
499
499
10 s
1000 Hz
12-lead
5,988
PTB-XL
18,885
21,837
— (unlabeled)
500 Hz
12-lead
262,044
Evaluation
LUDB
200
200
10 s
500 Hz
12-lead
2,369
Zhejiang
334
334
1.3–7.1 s
2000 Hz
12-lead
4,008
RDB
2,399
2,399
10 s
500 Hz
12-lead
28,788
TABLE I: Characteristics of the ECG datasets used for developing and evaluating the delineation model. Lead configurations range from 12-lead (six limb and six precordial leads) to 6-lead (limb only) and 2-lead (arbitrary pairs, e.g., MLII and V1).
mIoU (%) ↑
Summary
Pretraining
Fine-tuning
Internal
LUDB
Zhejiang
RDB
Avg. (%) ↑
Rank ↓
–
Supervised
75.2±0.2
64.8±0.8
53.3±2.4
60.5±0.5
63.5
3.50
Semi-supervised
81.6±0.2
70.0±5.9
66.0±5.3
67.9±0.9
71.4
MAE ∗
Supervised
82.8±0.3
71.3±2.3
75.2±1.5
71.9±1.2
75.3
1.25
Semi-supervised
82.3±0.5
72.8±5.3
74.8±2.1
71.1±0.4
75.2
MoCo ∗
Supervised
82.2±0.1
69.5±3.3
66.2±3.7
71.6±1.5
72.4
1.75
TABLE II: Benchmarking results of learning strategies for ECG delineation. Per-dataset entries are mIoU, reported as mean ± standard deviation over three seeds, where Internal denotes the QTDB + ISP test split. Under Summary, Avg. is the mIoU averaged across the four datasets for each fine-tuning strategy, and Rank is the average rank among the four pretraining strategies, aggregated over fine-tuning and datasets. The best value in each column is shown in bold .
mIoU (%) ↑
Averaged interval error (ms) ↓
Averaged point-wise sensitivity (%) ↑
Method
Internal
LUDB
Zhejiang
RDB
Internal
LUDB
Zhejiang
RDB
Internal
LUDB
Zhejiang
RDB
Avg. Rank ↓
NeuroKit2 (CWT)
41.3
44.9
7.3
40.0
70.9
108.3
72.8
30.2
59.4
59.9
15.7
54.9
5.8
NeuroKit2 (DWT)
43.4
45.6
8.6
38.6
43.6
45.6
52.7
40.5
60.3
62.2
17.8
58.5
5.2
ECGdeli
67.5
59.4
41.9
54.8
24.6
29.3
30.4
30.0
82.4
80.6
56.3
76.4
3.4
Prominence
60.4
63.9
42.7
51.7
28.1
22.3
35.9
23.0
78.1
82.9
54.6
73.7
3.6
DL delineator (sup)
83.0
71.6
73.5
71.1
11.5
19.0
16.3
22.2
96.1
86.6
82.9
92.6
1.8
TABLE III: Comparison of the DL delineator with open-source delineation tools. The best value in each column is in bold , and the second-best is underlined . Avg. Rank is the mean rank across all columns. CalECG , configured for interval outputs only, is compared separately on interval error in the Supplementary Material (Table S10 ).
Method
Average
PR
QRS
QT
NeuroKit2 (CWT)
97.3
207.4
32.1
52.3
NeuroKit2 (DWT)
51.3
45.3
17.0
91.5
ECGdeli
32.9
47.0
18.0
33.6
CalECG
37.9
41.1
32.7
40.0
Prominence
25.4
29.6
25.0
21.7
DL delineator (sup)
15.0
13.8
12.2
19.1
TABLE IV: Comparison of the DL delineator on mECGDB with open-source and commercial delineation tools, evaluated by mean absolute error (ms) for the PR interval, QRS duration, QT interval, and their average. The best value is in bold and the second-best is underlined .
Dataset
Method
P on
P off
QRS on
QRS off
T on
T off
Total
LUDB
NeuroKit2 (CWT)
43.7
43.1
80.8
72.2
56.4
63.4
59.9
NeuroKit2 (DWT)
73.3
63.8
66.2
81.8
32.5
55.8
62.2
ECGdeli
71.8
87.2
97.3
85.1
70.7
71.7
80.6
CalECG
74.0
-
67.7
69.1
-
61.5
68.1
Prominence
79.2
80.8
91.9
85.8
79.1
80.7
82.9
DL delineator (sup)
80.4
81.3
99.1
98.9
77.8
81.6
86.6
TABLE V: Comparison of the DL delineator on external test sets ( LUDB and Zhejiang ) with open-source and commercial delineation tools, evaluated by point-wise sensitivity (%) per fiducial point. ”Total” is the average over the available onset/offset points. The best value is in bold and the second-best is underlined .
Subgroup performance
Metric
Method
Sinus
Arrhythmia
Degradation
mIoU (%) ↑
ECGdeli
66.3
38.7
41.6%
Prominence
64.4
33.2
48.5%
DL delineator (sup)
76.1
64.9
14.7%
DL delineator (semi)
76.6
64.2
16.2%
Averaged interval error (ms) ↓
ECGdeli
28.9
40.1
38.8%
TABLE VI: Rhythm-stratified subgroup performance on RDB for the sinus and arrhythmia groups, with the relative performance degradation from sinus to arrhythmia (smaller is better). The best value is in bold and the second-best is underlined .
Fig. 2: Qualitative examples of ECG delineation results across methods and datasets. The rhythm classes shown reflect the characteristics of each dataset. Rows are grouped into the reference annotation (GT), rule-based tools, and the DL delineator ; shaded bands mark each method’s detected P (blue), QRS (orange), and T (green) waves on the same input signal. SR, sinus rhythm; PVC, premature ventricular contraction; AFIB, atrial fibrillation; SVT, supraventricular tachycardia.
Source
Train
Validation
Test
Total
QTDB
422
148
148
718
ISP
3,792
1,272
924
5,988
Internal ( QTDB + ISP )
4,214
1,420
1,072
6,706
TABLE S1: The detail number of the Internal ( QTDB + ISP ) set sample.
Dataset
Rhythm
#ECGs
LUDB
Sinus rhythm
143
Sinus tachycardia
4
Sinus bradycardia
25
Sinus arrhythmia
8
Irregular sinus rhythm
2
Abnormal rhythm (AFib, flutter)
18
TABLE S2: Rhythm composition of the LUDB and Zhejiang used for external test.
Group
Rhythm
#ECGs
Sinus
Sinus rhythm
400
Sinus bradycardia
400
Sinus tachycardia
140
Sinus irregularity
399
Arrhythmia
Atrial flutter
400
Atrial fibrillation
400
TABLE S3: Rhythm composition of the RDB used for the rhythm-stratified subgroup analysis. Recordings are grouped into sinus and arrhythmia categories.
Component
Setting
Framework
Mean Teacher (MT)
Teacher EMA decay
0.99
Consistency loss weight
1.0
Weak augmentation
Random resized cropping
Strong augmentation
RandAugment( N=3,p=0.5 )
Candidate transforms
Powerline noise
TABLE S4: Semi-supervised learning configuration.
Hyperparameter
MAE
MoCo
MERL
Epochs
800
300
50
Warm-up epochs
40
40
5
Batch size
2048
2048
2048
Learning rate
1.2e-3
1.2e-3
2.0e-4
Weight decay
0.05
0.1
1e-5
(β1,β2)
(0.9, 0.95)
(0.9, 0.999)
(0.9, 0.999)
TABLE S5: Optimization hyperparameters for self-supervised pretraining: MAE , MoCo , and MERL . All algorithms are trained with the AdamW optimizer and cosine learning rate scheduling.
Hyperparameter
Value
MAE
Masking ratio
0.75
Decoder embedding dimension
96
Decoder depth
4
Decoder attention heads
3
MoCo
TABLE S6: Model architecture and algorithm-specific hyperparameters for self-supervised pretraining.
Method
Time (ms) ↓
NeuroKit2 (CWT)
56.3 ± 11.9
NeuroKit2 (DWT)
63.6 ± 11.1
ECGdeli
141.0 ± 57.5
Prominence
20.0 ± 6.0
DL delineator
17.1 ± 1.9
TABLE S7: Computation time analysis on LUDB , reported as a mean ± standard deviation. The best value is bolded.
Pretraining
Fine-tuning
Internal
LUDB
Zhejiang
mECGDB
RDB
–
Supervised
16.2±0.3
21.1±0.7
33.1±2.0
26.1±0.9
23.4±0.3
Semi-supervised
12.8±0.4
20.8±2.4
21.7±3.0
16.2±0.6
29.8±4.3
MAE
Supervised
11.6±0.2
19.2±1.7
16.5±1.2
15.0±0.2
20.9±1.2
Semi-supervised
11.7±0.3
18.4±2.0
15.9±0.6
15.7±0.7
22.7±1.0
MoCo
Supervised
12.0±0.3
21.3±2.5
21.3±1.0
15.2±0.8
22.4±1.2
Semi-supervised
11.8±0.3
17.2±0.8
16.9±0.8
15.4±0.7
22.3±0.5
TABLE S8: Benchmarking results of learning strategies for ECG delineation. Per-dataset entries are averaged interval error (ms), reported as mean ± std over three seeds, where Internal denotes the QTDB + ISP test split. The best value in each column is shown in bold .
Pretraining
Fine-tuning
Internal
LUDB
Zhejiang
RDB
–
Supervised
91.3±0.2
82.5±1.2
66.5±2.6
87.1±0.7
Semi-supervised
95.4±0.2
86.5±7.0
76.6±6.2
90.8±0.6
MAE
Supervised
96.1±0.2
86.4±3.0
85.4±2.2
93.3±0.6
Semi-supervised
96.2±0.3
89.2±3.3
85.9±1.5
92.6±0.3
MoCo
Supervised
95.7±0.1
85.1±3.5
77.5±4.2
92.4±1.0
Semi-supervised
96.1±0.2
90.3±3.5
84.5±0.7
92.9±1.0
TABLE S9: Benchmarking results of learning strategies for ECG delineation. Per-dataset entries are averaged point-wise sensitivity (%), reported as mean ± std over three seeds, where Internal denotes the QTDB + ISP test split. The best value in each column is shown in bold .
Averaged interval error (ms) ↓
Method
Internal
LUDB
Zhejiang
CalECG
37.9
20.1
58.2
DL delineator (sup)
11.5
19.0
16.3
DL delineator (semi)
12.0
16.8
15.3
TABLE S10: Comparison of the DL delineator with CalECG on averaged interval error. The best value is in bold and the second best is underlined .
Sensitivity (%)
PPV (%)
Dataset
Method
40 ms
150 ms
40 ms
150 ms
Internal
NeuroKit2 (CWT)
59.4
73.7
66.1
82.1
NeuroKit2 (DWT)
60.3
88.4
60.5
88.0
ECGdeli
82.4
97.0
77.3
90.8
Prominence
78.1
95.9
74.2
91.1
DL delineator (sup)
96.1
99.0
95.6
98.6
TABLE S11: Comparison of the DL delineator with open-source delineation tools, evaluated by averaged point-wise total sensitivity and positive predictive value (PPV) (%) at tolerances of 40 and 150 ms. The best value in each column is in bold .
Dataset
Method
P on
P off
QRS on
QRS off
T on
T off
Total
LUDB
NeuroKit2 (CWT)
43.5
42.8
87.4
93.8
67.5
76.0
68.5
NeuroKit2 (DWT)
62.4
53.6
75.1
92.3
32.7
56.5
62.1
ECGdeli
55.1
66.9
96.8
84.6
63.2
64.1
71.8
CalECG
83.8
-
95.0
96.9
-
84.5
90.1
Prominence
64.6
65.9
91.0
84.9
79.0
80.7
77.7
DL delineator (sup)
88.1
89.1
99.2
99.0
81.4
85.4
90.4
TABLE S12: Comparison of the DL delineator on external test sets ( LUDB and Zhejiang ) with open-source and commercial delineation tools, evaluated by point-wise positive predictive value (PPV, %) per fiducial point. ”Total” is the average over the available onset/offset points. The best value is in bold .
Dataset
Method
P on
P off
QRS on
QRS off
T on
T off
LUDB
NeuroKit2 (CWT)
-7.6 ± 9.4
0.1 ± 9.5
-8.8 ± 12.3
-12.0 ± 13.3
5.0 ± 17.4
6.1 ± 10.9
NeuroKit2 (DWT)
3.5 ± 17.6
-3.0 ± 17.3
-3.6 ± 13.7
-3.1 ± 13.4
17.7 ± 20.1
-10.0 ± 13.2
ECGdeli
-13.0 ± 11.3
16.8 ± 11.4
-7.2 ± 13.9
16.2 ± 16.0
-10.4 ± 19.9
4.5 ± 15.9
CalECG
0.3 ± 11.5
-
8.1 ± 13.5
-2.3 ± 11.8
-
5.3 ± 12.1
Prominence
1.9 ± 12.8
3.1 ± 13.9
-0.0 ± 15.9
-11.4 ± 18.1
2.8 ± 18.2
13.2 ± 14.8
DL delineator (sup)
-5.4 ± 14.3
7.6 ± 15.1
-8.4 ± 10.1
10.4 ± 11.4
4.0 ± 20.1
12.5 ± 14.1
TABLE S13: Comparison of the DL delineator on LUDB and Zhejiang external test sets with widely used open-source and commercial delineation tools, evaluated by point-wise localization error (ms) per fiducial point at the 40 ms tolerance. The signed mean error and its standard deviation are reported.
Sinus group
Arrhythmia group
Metric
Method
SR
SB
ST
SI
Avg.
AF
AFIB
AT
SVT
Avg.
Deg.%
mIoU (%) ↑
ECGdeli
67.8
67.2
63.2
66.9
66.3
38.8
42.1
45.3
28.8
38.7
41.5
Prominence
67.2
65.4
58.9
66.3
64.4
33.1
38.1
38.6
22.8
33.2
48.6
DL delineator (sup)
76.4
77.0
75.0
75.9
76.1
61.1
67.6
63.5
67.4
64.9
14.7
DL delineator (semi)
77.3
78.0
74.2
76.9
76.6
59.4
67.8
63.4
66.3
64.2
16.2
Averaged interval error (ms) ↓
ECGdeli
28.3
30.1
28.4
28.6
28.9
34.1
29.8
34.4
62.0
40.1
65.8
TABLE S14: Rhythm-stratified subgroup analysis on the RDB , comparing the DL delineator with the two strongest tools (Prominence, ECGdeli). Rhythms are grouped into sinus (SR, SB, ST, SI) and arrhythmia (AF, AFIB, AT, SVT). For each metric, we report the group mean performance and its relative degradation (Deg.%) from sinus to arrhythmia. The best value is in bold .
Fig. S1: Qualitative P/QRS/T delineation on a representative 2.5 s Lead II segment from LUDB . Rows are grouped into the reference annotation (GT), rule-based tools, and the DL delineator ; shaded bands mark each method’s detected P (blue), QRS (orange), and T (green) waves on the same input signal.
Fig. S2: Qualitative P/QRS/T delineation on a representative 2.5 s Lead II segment from Zhejiang . Rows are grouped into the reference annotation (GT), rule-based tools, and the DL delineator ; shaded bands mark each method’s detected P (blue), QRS (orange), and T (green) waves on the same input signal.
Fig. S3: Qualitative P/QRS/T delineation on a representative 2.5 s Lead II segment from RDB . Rows are grouped into the reference annotation (GT), rule-based tools, and the DL delineator ; shaded bands mark each method’s detected P (blue), QRS (orange), and T (green) waves on the same input signal. SR, sinus rhythm; AFIB, atrial fibrillation; SVT, supraventricular tachycardia.
Department of Computer Science and Engineering, Indian Institute of Technology (ISM) Dhanbad, Dhanbad, Jharkhand, India. · Department of Computer Science and Engineering, Indian Institute of Technology Roorkee, Roorkee, Uttarakhand, India.
Department of Computer Engineering, University of Kurdistan, Sanandaj, Iran · Department of Mathematics and Operational Research, University of Mons, Mons, Belgium