NeuroECG: ECGFounder-Based Deep ECG Representation for EEG-Free Neurological Prognostication After Cardiac Arrest
Authors: Jiajun Gao, Yi Zhao, Chenyang Xu, Yuxi Zhou, Hao Wang
Organizations: School of Economics and Management Xidian University Xi’an, China · School of Mechano-Electronic Engineering Xidian University Xi’an, China · School of Cyber Engineering Xidian University Xi’an, China · DCST, BNRist, RIIT, Institute of Internet Industry Tsinghua University Beijing, China
Neurological prognostication after cardiac arrest commonly relies on electroencephalography (EEG). However, EEG demands high clinical resources. Bedside electrocardiography (ECG) is standard and low-cost. Yet, its value for predicting neurological outcomes remains underexplored. In this study, we propose NeuroECG, an ECGFounder-based deep representation framework for EEG-free auxiliary prognostication. NeuroECG adapts a pretrained ECG foundation model via task-specific fine-tuning. We implement a gradual unfreezing strategy on single-channel bedside monitoring ECG. Multiple ECG segments per patient are encoded into segment-level deep features. These embeddings are aggregated via quantile pooling (q = 0.24) and compressed using principal component analysis (PCA). Experiments on 412 ECG-available patients from the multicenter I-CARE database show that the adapted ECGFounder backbone achieves the best performance among ECG-only backbone baselines, with a test AUROC of 0.7333. We further combine the learned deep ECG representation with static clinical covariates. The proposed NeuroECG model achieves a test AUROC of 0.8077 and an AUPRC of 0.8970. These results support deep bedside ECG representations as a useful source of auxiliary prognostic information. Their integration with static clinical covariates improves prediction in an EEG-free setting. The source code is available at https://github.com/goddream66/NeuroECG
Figures & tables
Fig. 1: Overview of the proposed ECGFounder-based neurological prognostication framework. Bedside monitoring ECG recordings within the first 72 hours after ROSC are divided into 10-second single-channel segments and encoded by the ECGFounder backbone. Segment-level deep embeddings are aggregated into a patient-level representation using quantile pooling and further compressed by PCA. The compact deep ECG representation is combined with static clinical covariates for patient-level CatBoost prediction.
Backbone
Deep-only (Raw)
Deep PCA-only
test AUROC
test AUPRC
test CPC-MAE
test AUROC
test AUPRC
test CPC-MAE
ConvNeXt1D
0.5714±0.0200
0.7382±0.0238
1.6865±0.0105
0.4884±0.0387
0.6734±0.0378
1.7156±0.0097
SE-ResNet
0.5664±0.0228
0.7491±0.0231
1.6804±0.0177
0.5215±0.0573
0.6987±0.0469
1.6952±0.0125
ECG-JEPA †
0.5828±0.0426
0.7567±0.0323
1.6787±0.0125
0.6510±0.0911
0.7857±0.0637
1.6832±0.0272
ECG-FM †
0.6649±0.0365
0.8249±0.0292
1.6512±0.0213
0.6950±0.0136
0.7976±0.0123
1.6254±0.0232
ST-MEM †
0.5659±0.0763
0.7320±0.0651
1.7045±0.0097
0.6172±0.0353
0.7666±0.0299
1.6704±0.0125
TABLE I: Performance comparison of different ECG backbones under the same ECG-only and PCA-compressed evaluation settings. † denotes backbones initialized with pretrained weights.
Backbone
Deep PCA + Static
test AUROC
test AUPRC
test F1-score
test Recall
test CPC-MAE
SE-ResNet
0.7134±0.0220
0.8014±0.0195
0.7989±0.0242
0.8143±0.0315
1.6204±0.0381
ConvNeXt1D
0.7356±0.0225
0.8464±0.0147
0.7805±0.0396
0.7810±0.0259
1.6289±0.0383
ECG-FM †
0.7655±0.0199
0.8601±0.0213
0.8037±0.0153
0.8095±0.0240
1.5508±0.0553
ECG-JEPA †
0.7660±0.0121
0.8619±0.0073
0.7960±0.0171
0.8190±0.0358
1.5923±0.0515
ST-MEM †
0.7680±0.0139
0.8758±0.0167
0.8175±0.0238
0.8333±0.0336
1.5537±0.0552
TABLE II: Downstream outcome classification (AUROC, AUPRC, F1-score, Recall) and cerebral performance category regression (CPC-MAE) results of different backbones under the proposed NeuroECG configuration (Deep PCA + Static).
Features
Dim.
AUROC
AUPRC
CPC-MAE
Static only
6
0.7689±0.0034
0.8574±0.0023
1.5762±0.0274
Deep only
1024
0.7333±0.0187
0.8540±0.0163
1.5655±0.0172
Deep PCA-only
64
0.7097±0.0108
0.8011±0.0231
1.5890±0.0202
Static + Deep
1030
0.7931±0.0040
0.8866±0.0033
1.5494±0.0026
Static + Deep PCA
70
0.8077±0.0170
0.8970±0.0082
1.5156±0.0364
TABLE III: Feature ablation on the test set (mean ± standard deviation).
Configuration
AUROC
AUPRC
Mean pooling
0.7730±0.0233
0.8615±0.0285
Max pooling
0.7662±0.0098
0.8607±0.0099
Quantile pooling ( q=0.24 )
0.8077±0.0170
0.8970±0.0082
Head-only fine-tuning
0.6968±0.0181
0.8499±0.0125
Immediate full unfreezing
0.6213±0.0146
0.7759±0.0135
Gradual unfreezing
0.7333±0.0187
0.8540±0.0163
TABLE IV: Key pooling and fine-tuning ablations on the test set. Pooling uses Static + Deep PCA; fine-tuning uses ECG-only features.
Fig. 2: Test-set SHAP summary for the fused NeuroECG model. The 15 leading features include static clinical variables and PCA-compressed ECG components. Positive SHAP values indicate a larger predicted probability of poor neurological outcome.
Myocardial infarction (MI) is a leading cause of death, and its adverse outcomes are urgent to predict. Yet ECG-based prognostic models underperform because deep learning requires large, labelled datasets, which are scarce in medicine. Foundation models can learn from unlabelled ECGs via selfsupervision, but medically relevant training strategies remain underexplored. We propose a pretrained artificial intelligence model that combines patient-specific temporal information using contrastive learning with supervised multitask heads, then fine-tunes on post-MI outcome prediction. The proposed model outperformed a model trained from scratch (0.794 vs 0.608 AUC) showing that clinically structured ECG modelling improves classification in limited data regimes.
Electrocardiogram (ECG) interpretation requires knowledge of cardiology, electrophysiology, clinical diagnosis, ECG waveforms, signal acquisition, and instrumentation. Existing language-model benchmarks, however, primarily assess broad medical knowledge or interpretation of individual ECG signals and images rather than the broader contextual knowledge required for ECG interpretation. We developed ECGQuest, a literature-grounded resource for evaluating and fine-tuning ECG-specific language models. A GPT-4o-based pipeline generated questions from 23 ECG references and Computing in Cardiology proceedings from 2003-2025. The final dataset contains 10,904 unique True/False questions paired with their negated forms (21,808 Q&A pairs). We evaluated three commercial and 20 open-source language models on a held-out test set in a zero-shot setting. Five open-source models with 7-14B parameters were fine-tuned using Low-Rank Adaptation, with BERT and BiomedBERT included as supervised encoder baselines. Generalization was assessed on ECG-related subsets of MedMCQA and MedQA converted to binary True/False questions using official answer keys. Zero-shot accuracy on ECGQuest ranged from 49.5% to 74.4%, with GPT-5 performing best. General-purpose models outperformed medically specialized models, several models showed strong True/False bias, and encoder baselines performed near chance. Fine-tuning improved all open-source models by 6.5-14.1%. Fine-tuned DeepSeek-R1-Distill-Qwen-14B reached 76.3% accuracy, while a five-model voting ensemble reached 78.5%. On MedMCQA and MedQA, fine-tuning mainly benefited weaker or class-biased models and did not consistently improve strong base models. ECGQuest provides a reproducible benchmark for contextual ECG knowledge and shows that parameter-efficient fine-tuning can make smaller language models competitive with substantially larger commercial models.
Mohammadsina Hassannia, Matthew A. Reyna, Reza Sameni
Department of Biomedical Informatics Emory University Atlanta, GA 30322, USA · Department of Biomedical Engineering Georgia Institute of Technology and Department of Biomedical Informatics Emory University Atlanta, GA 30332, USA
Electrocardiograms (ECGs) are widely used non-invasive measurements of cardiac activity and play a central role in clinical diagnosis. Recent multimodal approaches align ECG signals with clinical reports to incorporate diagnostic semantics, but clinical reports often fail to preserve the rich physiological structure of ECG waveforms, particularly across multiple levels of abstraction ranging from coarse diagnostic categories to fine-grained morphology. To address this limitation, we formulate ECG representation learning from an information-theoretic perspective and derive a tractable objective that jointly preserves signal structure and integrates clinical semantics. Based on this principle, we propose \textbf{MERIT} (Multimodal ECG Representation via Information Theory), a dual-branch pretraining framework combining masked ECG modeling with ECG--text contrastive alignment. Extensive experiments on PTB-XL and additional benchmarks demonstrate consistent improvements over prior methods, including gains exceeding 3 F1 on PTB-XL All and 5 F1 on SubClass classification. In zero-shot evaluation, MERIT further improves performance by up to +2.66% AUC and +2.11% F1 on PTB-XL SubClass, while also demonstrating robustness under multiple distribution-shift settings. Moreover, leveraging the learned ECG representations for ECG-conditioned clinical text generation with large language models improves text quality across several metrics, including ROUGE and METEOR. Together, these results demonstrate that MERIT learns more informative and clinically meaningful ECG representations, particularly for fine-grained clinical applications.