Teaching PPG How not Who: Fixed-Effects Distillation from ECG
Authors: Zhongli Wu, Zhuangzhi Gao, Yuankai Wang, Gregory Y. H. Lip, Bilal H. Kirmani, Yalin Zheng
Organizations: Liverpool Centre for Cardiovascular Science, University of Liverpool, UK · Shanghai Artificial Intelligence Laboratory, Shanghai, China · Liverpool Heart and Chest Hospital, Liverpool, UK · Department of Eye and Vision Sciences, University of Liverpool, UK
ECG is widely used to teach PPG-only models, yet what it teaches is unexamined. Wearables are valued for tracking how a person's cardiovascular state changes, but ECG-to-PPG distillation mostly learns who the person is. A per-recording mean, the trait, holds 40-59% of a frozen ECG teacher's target, and pooled students memorise it without carrying it to new recordings. The raw alignment cosine misses this, since a constant predictor scores 0.793. Across 34 runs, the more identity a student memorises, the less state it learns. Fixed-effects distillation subtracts each recording's mean from prediction and target, so the trait cancels exactly, while a pooled anchor keeps it. State agreement more than doubles, within-person labels improve while age and sex do not, and the gain holds on two backbones and two further databases. Conditioning on the recording turns distillation toward the within-person changes that wearables monitor.
Figures & tables
Figure 1: (a) Training. The frozen ECG teacher T is split after its fourth stage: its lower layers Tlo map ECG to 400 -channel tokens hri , and the bridge fθ predicts these tokens as h^ri from multi-scale features of the frozen PPG backbone. The upper layers Thi map hri and h^ri to zri and z^ri . The anchor compares raw embeddings, and with equal weight the logits of the teacher’s frozen classification head, so it keeps the trait μr ; the state loss compares deviations from the mean over the windows of recording r in the batch, where μr cancels, with candidates Cr from the same recording. (b) Inference needs PPG only (about 39 M parameters with PaPaGei-S, 28.6 M of them in Thi ); a linear read-out fitted once maps z^ri to downstream tasks.
Figure 2: Teacher targets of four training recordings (360 windows each); no model is involved. Each panel uses its own principal axes, both at one scale, and each ellipse covers two standard deviations of one recording. Left: the anchor compares raw targets, where the row means μr ( × ) set the recordings apart; the trait holds 59% of the centred energy, here as over all training recordings. Right: the state loss compares eri , where every row mean moves to the origin and only the shape of each recording’s own variation remains. The pink recording visits two states, and its μr lies between them.
Figure 3: Colour marks the objective and marker shape the backbone in every panel. (a) Every run: trait agreement on training windows, which gauges memorised identity, against state agreement on held-out recordings; seed mean ± sd over three seeds where available (Sec. 4.1 ), and in grey the pooled loss on batches of 96×1 and 24×4 windows. (b) The same objectives on ECSMP, an external database with chest ECG and wrist PPG (seed mean ± sd). (c) The λ path of the linear analogue on ECSMP with one ridge penalty held fixed along the path, relative to the within estimator; both backbones decay monotonically towards the pooled solution.
state agreement ↑
trait agreement
raw cosine
pooled
0.136 ± 0.006
0.316 ± 0.003
0.798 ± 0.002
fe-nce
w/o anchor
0.210 ± 0.021
0.003 ± 0.002
0.109 ± 0.073
with anchor
0.268 ± 0.000
0.220 ± 0.031
0.797 ± 0.005
fe-cos
w/o anchor
0.249 ± 0.012
0.013 ± 0.007
0.174 ± 0.040
with anchor
0.307 ± 0.001
0.219 ± 0.026
0.795 ± 0.007
Table 1: Held-out agreement on PaPaGei-S without and with the anchor (three seeds).
PaPaGei-S
AnyPPG
none
pooled
mix-NCE
fe-cos
+fe-nce
pooled
mix-NCE
fe-cos
+fe-nce
held-out recordings: labels that vary within a recording
SBP error (mmHg)
10.97
8.93 ± 0.03
8.86 ± 0.02
8.77 ± 0.05
8.63 ± 0.06
8.21 ± 0.01
8.16 ± 0.09
7.83 ± 0.04
7.79 ± 0.03
DBP error (mmHg)
5.48
4.81 ± 0.02
4.80 ± 0.00
4.73 ± 0.02
4.69 ± 0.03
4.50 ± 0.02
4.45 ± 0.04
4.25 ± 0.02
4.22 ± 0.02
HR error (bpm)
5.90
2.50 ± 0.02
2.62 ± 0.02
2.29 ± 0.01
2.50 ± 0.01
2.67 ± 0.02
2.84 ± 0.03
2.38 ± 0.00
2.62 ± 0.02
state retrieval top-1
0.004
0.032 ± 0.001
0.123 ± 0.001
0.112 ± 0.001
0.196 ± 0.003
0.040 ± 0.001
0.138 ± 0.001
0.137 ± 0.001
0.221 ± 0.001
Table 2: Public labels and a second database. Linear read-outs of the predicted embedding, fitted on training recordings and scored on 75 held-out recordings never used for selection and on 60 MIMIC-III recordings the bridge never saw. Errors of labels that vary within a recording are taken after removing each person’s mean; ‘none’ is the no-information reference. Mean ± sd over three seeds; mix-NCE: InfoNCE with negatives from all recordings; best objective per backbone in bold.
Electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG) provide complementary views of the same cardiac cycle, yet existing cardiac foundation models are trained for a single sensing modality, leaving the shared physiology across sensors unexploited. We introduce CardioState-JEPA, a cardiac foundation model to learn a single shared representation jointly across ECG, PPG, and PCG, built on a physiology-aware joint-embedding predictive architecture. The model maps heterogeneous waveforms into a common token space, processes them with a single shared Transformer encoder, and learns by predicting masked latent cardiac states, placing the pretraining target on shared physiology rather than sensor-specific waveform appearance. To handle the temporal offsets between electrical, mechanical, and hemodynamic events, cross-modal prediction uses a learned delay aligner that matches signals at the corresponding cardiac time. Because synchronized multi-sensor recordings are scarce, CardioState-JEPA first learns within-modality structure from abundant unimodal data and then uses paired data to align modalities in latent cardiac time. Evaluated as a frozen encoder across 25 downstream tasks spanning ECG, PPG, and PCG, our encoder improves average PPG classification by 8.2 AUROC points, PCG murmur detection by 18.8 AUROC points, and ECG classification by 15.5 AUROC points over the best self-supervised signal baseline and matches or exceeds cardiac models trained with privileged clinical text or supervised labels on several ECG benchmarks. These results establish that heterogeneous cardiac signals can mutually supervise a single foundation model of cardiac physiology.
Hamza Shafiq, Hung Manh Pham, Bin Zhu +3
Eindhoven University of Technology, Netherlands · Singapore Management University, Singapore
Electrocardiography (ECG) is the clinical standard for cardiac assessment but requires dedicated hardware that does not scale to daily-life monitoring. Photoplethysmography (PPG) is ubiquitous in wearables but lacks ECG-specific diagnostic morphology and is corrupted by motion and sensor noise. PPG-to-ECG generation aims to bridge this gap by recovering electrical morphology and timing from peripheral pulse signals. However, existing methods largely rely on statistical alignment and data-driven generation. They fail to explicitly structure the latent space around physiology-aware electro-hemodynamic factors and lack constraints from forward physiological dynamics. To address these challenges, we propose PG-LRF, a physiology-guided latent rectified flow framework. PG-LRF introduces an electro-hemodynamic simulator that co-models ECG and PPG through shared cardiac phase dynamics. Guided by this simulator, a Physiology-Aware AutoEncoder learns a structured electro-hemodynamic latent space. Then we integrate this simulator guidance into a PPG-conditioned latent rectified flow, enforcing ECG-side morphology consistency and ECG-to-PPG forward hemodynamic consistency during generative transport. Experiments on the large-scale MC-MED dataset demonstrate that PG-LRF significantly improves PPG-to-ECG generation and downstream cardiovascular disease classification, proving its ability to generate ECGs that are both signal-faithful and physiologically plausible under the ECG-to-PPG hemodynamic pathway
Xiaoda Wang, Minxiao Wang, Kaiqiao Han +10
Department of Computer Science, Emory University · Department of Computer Science, University of California, Los Angeles · Nell Hodgson Woodruff School of Nursing, Emory University +2
High-fidelity ECG interpretation is increasingly reliant on massive foundation models, yet their deployment in clinical edge-care remains hindered by extreme computational demands. While knowledge distillation (KD) is a promising solution, traditional methods fail to capture the complex spatio-temporal dependencies of ECG signals when transferring knowledge across heterogeneous architectures. In this paper, we propose EVL-ECG, a framework specifically designed for cross-architecture distillation of cardiac diagnostic logic. EVL-ECG introduces three ECG-aware innovations: (1) Multi-Head Cross-Attention Alignment, which harmonizes architectural discrepancies to preserve fine-grained morphological features; (2) Optimal Transport-based Visual Feature Matching, utilizing optimal transport to maintain global structural relationships across ECG leads despite mismatched token representations; and (3) Geometric Intra-Architecture Relation Matching, which distills the latent diagnostic reasoning of the teacher model. Evaluations across ECG benchmarks demonstrate that EVL-ECG yields improvements of up to 2.4% AUC and 1.1% clinical accuracy over existing baselines. Notably, EVL-ECG establishes an efficient 2B-parameter ECG foundation model, suitable for resource-constrained clinical environments.
Hanoi University of Science and Technology, Hanoi, Vietnam · VinUni-Illinois Smart Health Center, VinUniversity, Hanoi, Vietnam · College of Engineering and Computer Science, VinUniversity, Hanoi, Vietnam +1