A Hybrid CNN--State-Space--Attention Backbone with Joint-Embedding Predictive Pretraining for 12-Lead ECG Classification
Authors: Yakoub Bazi, Sarah Aljuhani, Mohamad M. Al Rahhal, Mansour Zuair, Naif Alajlan
Organizations: Computer Engineering Department, College of Computer and Information Sciences, King Saud University, Riyadh, 11543, Saudi Arabia · Applied Computer Science Department, College of Applied Computer Science, King Saud University, Riyadh, 11543, Saudi Arabia
Automatic 12-lead electrocardiogram (ECG) classification requires representations that jointly capture local waveform morphology, long-range temporal dynamics, and cross-lead dependencies, yet integrating these properties within a single efficient architecture remains challenging. This paper introduces a hybrid CNN-SSM-Attention backbone for 12-lead ECG classification. A convolutional stem performs early waveform tokenization and temporal reduction, mixed state-space and depthwise-convolutional blocks model temporal dynamics and local morphology, and a late self-attention stage enables global token interaction at reduced resolution. To improve transfer from unlabeled data, we further develop an ECG-oriented Joint-Embedding Predictive Pretraining (JEPA) framework. Unlike ViT-based JEPA methods that mask patch tokens before the encoder, the proposed method samples span masks at the latent temporal resolution and projects them back to the waveform domain, then predicts clean latent targets from a momentum encoder without waveform reconstruction. Experiments on CPSC2018, Chapman-Shaoxing, and PTB-XL, with pretraining on approximately 350K unlabeled CODE-15 recordings, show that the proposed backbone provides strong supervised baselines under a compact parameter budget. JEPA pretraining further improves transfer, particularly in reduced-label settings and under both full fine-tuning and LoRA-based adaptation. Code: https://github.com/yakoubbazi/Hybrid_ECG_Jepa
Figures & tables
Figure 1 : Overview of the proposed hybrid CNN–SSM–Attention backbone for 12-lead ECG classification. The architecture combines early convolutional tokenization, hybrid mixer stages, late self-attention, and attention pooling for final classification.
Figure 2 : Architecture of a Stage 1 block. The block operates at resolution T1×C1 and combines an SSM branch for temporal sequence modeling with a local gated depthwise-convolution branch for morphology-aware refinement, followed by residual fusion and an MLP sublayer.
Figure 3 : Architecture of a Stage 3 block. Operating at resolution T3×C3 , each block follows a Transformer-style design composed of a multi-head self-attention sublayer and an MLP sublayer, with a residual connection around each component.
Variant
C1
C2
C3
N1
N2
N3
Heads
Params
Micro
32
64
64
1
1
1
2
0.21M
Mini
48
96
96
1
2
2
2
0.78M
Tiny
64
128
256
2
2
2
4
3.11M
Small
64
128
256
2
2
3
4
4.17M
Table 1 : Scalable backbone configurations used in this work. C1 , C2 , and C3 are the channel dimensions of Stages 1–3, and N1 , N2 , and N3 are the numbers of blocks in each stage. Heads is the number of attention heads in the Stage 3 self-attention blocks, and Params is the total number of parameters.
Figure 4 : Overview of the proposed ECG-JEPA pretraining scheme. A clean 12-lead ECG segment is processed by the momentum target encoder fξ to produce the final Stage 3 target representation Z3 . A corrupted view of the same signal is processed by the online encoder fθ , followed by the predictor gϕ , to predict masked final-stage latent targets. The loss is computed only over masked temporal positions.
Dataset
Classes
Train
Val
Test
Majority
Minority
Max/Min
CPSC2018
9
5,121
640
640
RBBB (1,227)
LBBB (143)
8.58
Chapman-Shaoxing
4
6,871
859
859
SB (3,085)
GSVT (573)
5.38
PTB-XL
5
12,957
1,637
1,650
NORM (7,243)
HYP (415)
17.45
Table 2 : Downstream dataset statistics after single-label filtering. Majority/minority classes and Max/Min ratios are computed from the training split.
Variant
CPSC2018
Chapman-Shaoxing
PTB-XL
AUC
Acc
F1
κ
AUC
Acc
F1
κ
AUC
Acc
F1
κ
Micro
94.66
76.88
71.58
72.90
98.60
93.60
91.71
90.64
90.02
76.97
62.13
62.99
Mini
94.68
80.16
75.25
76.74
98.52
94.30
92.38
91.67
89.87
76.24
62.80
61.53
Tiny
94.95
78.91
73.23
75.25
98.10
94.06
90.70
91.26
89.14
74.42
61.06
58.57
Small
95.41
80.16
74.56
76.76
98.60
94.88
92.70
92.47
88.91
77.52
64.75
64.13
Table 3 : Scratch performance of the backbone variants on CPSC2018, Chapman-Shaoxing, and PTB-XL. All values are reported in percentage (%).
Figure 5 : Representative ECG waveform corruption under different temporal masking ratios. Temporal masks are sampled at the latent token resolution and projected back to the waveform domain before the online encoder. Lower masking preserves more visible waveform morphology, whereas higher masking imposes a stronger contextual prediction task. Lead dropout is applied as an additional channel-level corruption.
Mode
CPSC2018
Chapman-Shaoxing
PTB-XL
AUC
Acc
F1
κ
AUC
Acc
F1
κ
AUC
Acc
F1
κ
Scratch
95.41
80.16
74.56
76.76
98.60
94.88
92.70
92.47
88.91
77.52
64.75
64.13
SSL-LoRA
95.39
82.03
77.58
78.93
99.15
97.09
94.92
95.71
90.21
77.03
61.56
63.05
SSL-Full
96.67
82.97
78.58
80.03
98.14
96.27
94.10
94.55
89.92
77.82
61.95
63.08
Table 4 : Transfer results of the Small backbone on CPSC2018, Chapman-Shaoxing, and PTB-XL. SSL-pretrained models are adapted using LoRA and full fine-tuning. All values are reported in percentage (%).
Train Fraction
Mode
AUC
Acc
Macro-F1
κ
1%
Scratch
71.62 ± 1.24
36.98 ± 2.13
28.11 ± 2.90
26.35 ± 2.68
SSL-LoRA
81.37 ± 1.14
51.15 ± 3.56
39.58 ± 6.31
42.43 ± 4.17
SSL-Full
82.92 ± 1.00
54.53 ± 3.17
45.70 ± 4.33
46.60 ± 3.74
5%
Scratch
84.41 ± 0.52
57.03 ± 1.84
49.50 ± 2.73
49.84 ± 2.19
SSL-LoRA
92.06 ± 0.27
72.45 ± 0.72
65.30 ± 2.40
67.71 ± 1.00
SSL-Full
91.65 ± 0.50
71.82 ± 0.89
65.12 ± 2.08
66.99 ± 1.15
Table 5 : Low-label transfer results on CPSC2018 using different fractions of the labeled training data. SSL-pretrained models are adapted using LoRA and full fine-tuning. Results are reported as mean ± standard deviation over three runs, in percentage (%).
Train Fraction
Mode
AUC
Acc
Macro-F1
κ
1%
Scratch
85.05 ± 1.56
66.67 ± 1.34
59.20 ± 1.86
50.03 ± 2.04
SSL-LoRA
98.27 ± 0.17
90.84 ± 0.87
86.99 ± 1.72
86.49 ± 1.33
SSL-Full
98.15 ± 0.29
90.65 ± 1.05
86.61 ± 1.40
86.22 ± 1.57
5%
Scratch
92.51 ± 0.95
80.75 ± 3.54
76.09 ± 4.02
71.42 ± 5.78
SSL-LoRA
98.30 ± 0.22
91.73 ± 0.51
88.41 ± 0.69
87.92 ± 0.75
SSL-Full
98.37 ± 0.18
91.93 ± 0.64
88.78 ± 0.96
88.24 ± 0.95
Table 6 : Low-label transfer results on Chapman-Shaoxing using different fractions of the labeled training data. SSL-pretrained models are adapted using LoRA and full fine-tuning. Results are reported as mean ± standard deviation over three runs, in percentage (%).
Train Fraction
Mode
AUC
Acc
Macro-F1
κ
1%
Scratch
74.51 ± 0.81
62.53 ± 0.73
43.04 ± 1.27
39.77 ± 0.67
SSL-LoRA
75.88 ± 0.80
64.83 ± 2.43
43.36 ± 1.15
41.40 ± 2.41
SSL-Full
76.51 ± 1.51
65.27 ± 2.43
45.45 ± 0.65
43.29 ± 1.89
5%
Scratch
81.65 ± 1.19
70.08 ± 1.91
50.96 ± 0.83
50.74 ± 2.49
SSL-LoRA
82.51 ± 0.57
70.79 ± 1.36
52.11 ± 1.46
51.78 ± 1.38
SSL-Full
82.78 ± 1.32
70.46 ± 1.36
52.64 ± 1.52
52.06 ± 1.09
Table 7 : Low-label transfer results on PTB-XL using different fractions of the labeled training data. SSL-pretrained models are adapted using LoRA and full fine-tuning. Results are reported as mean ± standard deviation over three runs, in percentage (%).
Mask Ratio
CPSC2018
Chapman-Shaoxing
PTB-XL
AUC
Acc
F1
κ
AUC
Acc
F1
κ
AUC
Acc
F1
κ
25%
94.32
76.88
71.13
72.90
97.61
96.16
93.69
94.37
89.50
77.64
63.25
63.58
50%
96.30
83.13
79.19
80.26
98.57
96.16
94.29
94.35
89.42
76.73
63.69
62.73
75% (default)
96.67
82.97
78.58
80.03
98.14
96.27
94.10
94.55
89.92
77.82
61.95
63.08
Table 8 : Sensitivity of JEPA-style pretraining to the temporal masking ratio using the Small backbone and full fine-tuning. The 75% setting corresponds to the default configuration used in the main experiments. All values are reported in percentage (%).
Figure 6 : Qualitative interpretation of the Small backbone. Panels (a) and (b) show correct and incorrect predictions on PTB-XL, respectively, while panel (c) shows a misclassification on CPSC2018. In each panel, the upper strip presents the Grad-CAM++ map [ 4 ] computed from the Stage 2 representation, and the lower strip presents the attention-rollout map [ 1 ] aggregated from the Stage 3 self-attention blocks.
Method
Setting
CPSC2018
PTB-XL
AUC
Acc
F1
κ
AUC
Acc
F1
κ
Ma et al. (2026)
Random Init.
86.95
68.03
58.21
61.53
85.76
72.37
54.31
53.50
Ma et al. (2026)
MoCo
80.30
51.04
41.50
41.75
79.70
65.93
45.21
43.25
Ma et al. (2026)
NNCLR
86.86
64.36
54.98
58.00
85.76
74.39
55.75
56.38
Ma et al. (2026)
SimCLR
89.52
68.84
60.82
63.32
86.64
73.97
56.07
56.13
Ma et al. (2026)
DCCLR
88.40
67.39
59.57
61.61
87.21
74.24
56.10
56.67
Table 9 : Comparison at 20% labeled training data on CPSC2018 and PTB-XL. All values are reported in percentage (%). Bold values indicate the best result within each dataset and metric.
Method
Setting
CPSC2018
PTB-XL
AUC
Acc
F1
κ
AUC
Acc
F1
κ
Liu et al. (2023)
Frozen
92.01
68.07
63.22
–
86.76
76.00
57.27
–
Shi et al. (2024)
SSL-Full
95.38
74.40
–
–
91.23
78.87
–
–
Liu et al. (2025)
SSL-Full
93.98
–
68.82
–
90.46
–
61.69
–
Ma et al. (2026)
Proposed
92.53
77.47
69.99
–
90.23
76.86
63.11
–
Ours
Random Init.
95.41
80.16
74.56
76.76
88.91
77.52
64.75
64.13
Table 10 : Full-label results and comparison with recent ECG classification methods on CPSC2018 and PTB-XL. All values are reported in percentage (%).
Electrocardiogram (ECG) arrhythmia classification remains challenging due to signal variability, noise, limited labeled data, and the difficulty in achieving both accuracy and efficiency in models. While self-supervised learning reduces label dependency, most methods target either global contextual features or local morphological patterns, but rarely implement hierarchical multi-scale feature extraction. ECG signals require architectures that simultaneously capture fine-grained beat-level morphology and broader rhythm-level dependencies with computational efficiency. To overcome this limitation, this paper proposes the Electrocardiogram Neighborhood Attention Transformer (ECG-NAT), a novel self-supervised learning approach tailored for multi-lead ECG classification. Our two-stage approach begins with generative pretraining, using a masked autoencoder to reconstruct partially masked ECG signals across multiple diverse datasets, enabling the model to learn robust, domain-invariant representations from unlabeled data. This is followed by discriminative fine-tuning with a dual-loss function that combines supervised contrastive and cross-entropy losses, aligning representation learning with label prediction. The hierarchical attention mechanism efficiently captures multi-scale temporal features from localized beat morphology to broader rhythm patterns at low computational cost. ECG-NAT achieves robust performance on benchmark datasets, with 88.1% accuracy using only 1% labeled data, demonstrating strong efficacy in low-resource settings. The framework combines superior classification performance with computational efficiency, making it practical for real-time ECG diagnosis. The code will be made available upon acceptance at: https://github.com/Mahsagazeran/ECG-NAT.
Department of Computer Engineering, University of Kurdistan, Sanandaj, Iran · Department of Mathematics and Operational Research, University of Mons, Mons, Belgium
Data analysis in the medical domain often encounters scenarios involving a limited target dataset and a large, unannotated dataset with a general distribution. Under such circumstances, self-supervised learning (SSL) methods are highly effective for utilizing large datasets, making them a popular choice for electrocardiogram (ECG) analysis. This work presents the Event Reconstruction Joint-Embedding Predictive Architecture (ER-JEPA), a lightweight SSL framework for multivariate time series, whose name and two-fold hierarchical structure are inspired by the diagnostic approach of cardiologists. At its core, ER-JEPA features: (1) a two-stage structure that constructs representations for each time interval and subsequently processes these representations as a univariate time series, (2) the hierarchical integration of two Joint-Embedding Predictive Architectures (JEPAs), and (3) a Vision Transformer (ViT) backbone. The structural concatenation of two JEPAs categorizes the model as a Hierarchical JEPA (H-JEPA), designed to encode multiple levels of abstract representations for enhanced prediction on complex tasks. This study reports a successful application of H-JEPA to 12-lead ECG data as a multivariate time series, alongside an analysis of the sensitivity of hierarchical representation during the pretraining stage. Pretrained on approximately 180,000 10-second recordings, the model achieves state-of-the-art downstream performance on the ST-MEM benchmark, with rapid computation and minimal resource usage.
Siwon Kim
Research Institute of Basic Sciences, Seoul National University, Seoul, Korea
Electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG) provide complementary views of the same cardiac cycle, yet existing cardiac foundation models are trained for a single sensing modality, leaving the shared physiology across sensors unexploited. We introduce CardioState-JEPA, a cardiac foundation model to learn a single shared representation jointly across ECG, PPG, and PCG, built on a physiology-aware joint-embedding predictive architecture. The model maps heterogeneous waveforms into a common token space, processes them with a single shared Transformer encoder, and learns by predicting masked latent cardiac states, placing the pretraining target on shared physiology rather than sensor-specific waveform appearance. To handle the temporal offsets between electrical, mechanical, and hemodynamic events, cross-modal prediction uses a learned delay aligner that matches signals at the corresponding cardiac time. Because synchronized multi-sensor recordings are scarce, CardioState-JEPA first learns within-modality structure from abundant unimodal data and then uses paired data to align modalities in latent cardiac time. Evaluated as a frozen encoder across 25 downstream tasks spanning ECG, PPG, and PCG, our encoder improves average PPG classification by 8.2 AUROC points, PCG murmur detection by 18.8 AUROC points, and ECG classification by 15.5 AUROC points over the best self-supervised signal baseline and matches or exceeds cardiac models trained with privileged clinical text or supervised labels on several ECG benchmarks. These results establish that heterogeneous cardiac signals can mutually supervise a single foundation model of cardiac physiology.
Hamza Shafiq, Hung Manh Pham, Bin Zhu +3
Eindhoven University of Technology, Netherlands · Singapore Management University, Singapore