PPG-LM: A Photoplethysmography-Language Model with Multi-Level Clinical Alignment
Authors: Xiaoda Wang, Minxiao Wang, Maxwell A Xu, Patrick Langer, Kaiqiao Han, Defu Cao, Xiao Luo, Yuzhe Yang, +5 more
Organizations: Emory University · University of California, Los Angeles · Google · Stanford University · University of Southern California · University of Wisconsin–Madison
Photoplethysmography (PPG) is widely recorded by clinical monitors and consumer wearables, providing a scalable source of continuous physiological information. These recordings offer an opportunity for physiological assessment at scale, but realizing this potential requires models to learn from both signal-derived physiological supervision and broader clinical context captured in electronic health records (EHRs). This involves aligning information spanning local observations, care events, and entire visits with PPG representations at corresponding temporal scales. However, existing PPG foundation models primarily rely on task-specific prediction heads, while the medical knowledge of large language models does not necessarily translate into waveform understanding. To bridge this gap, we introduce PPG-LM, the first PPG-language model family to learn physiological representations from both signal-derived supervision and broader clinical context captured in EHRs. To construct clinically grounded captions, we develop an automatic captioning pipeline that generates segment-, event-, and visit-level descriptions from signal measurements and structured EHR records. We then learn from these pairs through a two-stage framework that first establishes segment-language correspondence through contrastive learning and waveform-conditioned captioning, then extends alignment to events and visits through time-aware aggregation and temporal statement matching. Pretrained on approximately 73k hours of PPG, PPG-LM supports language-based recognition, cross-modal retrieval, and segment captioning. Experiments on MC-MED, MIMIC-III, and VitalDB show improved retrieval and caption factuality over language-model baselines and gains over PPG and time-series foundation models on multiple clinical prediction tasks.
Figures & tables
Figure 1: Motivation and overview of PPG-LM. PPG FMs learn waveform representations with limited integration of language and EHR context. General LLMs capture medical knowledge from text but require adaptation to encode raw PPG. PPG-LM bridges this gap by pairing PPG with captions derived from physiological measurements and structured EHRs. The radar plot highlights PPG-LM’s superior performance over baselines on selected segment-, event-, and visit-level tasks.
Study
Signal
Text Source
Text scope
Raw
Context
Segment
Event
Visit
SensorLM ( Zhang et al., 2026a )
✗
1 day
Sensor
✓
✗
✗
NormWear ( Luo et al., 2024 )
✓
6 s
Labels
✓
✗
✗
CSFM ( Gu et al., 2026 )
✓
10 s
Labels
✓
✗
✗
PulseLM ( Pham et al., 2026 )
✓
10 s
Labels
✓
✗
✗
CAP ( He et al., 2026 )
✓
5 min
EHR
✓
✗
✗
Table 1: Comparison of PPG–language studies.
Figure 2: Clinically Grounded Caption Generation: (1) Segment-level captions describe physiological features and signal quality within 30-second PPG windows. (2) Event-level captions provide medication context for fixed windows before and after administration. (3) Visit-level captions associate variable-length recordings with demographics, and medical history from clinical EHRs.
Figure 3: Two-stage multilevel alignment. (a) Stage 1 learns PPG and text encoders through contrastive segment–language alignment and PPG-conditioned caption generation. (b) Stage 2 extends the shared encoders to event- and visit-level alignment.
Model
MC-MED
MIMIC-III
VitalDB
Rhythm ↑
Qual. ↑
Perf. ↑
Morph. ↑
RR ↑
Rhythm ↑
Qual. ↑
RR ↑
Rhythm ↑
Qual. ↑
RR ↑
GPT-5.6-luna
0.322
0.314
0.174
0.259
0.332
0.376
0.291
0.266
0.317
0.162
0.313
GPT-6-luna
0.319
0.361
0.193
0.263
0.259
0.357
0.315
0.253
0.326
0.175
0.307
GPT-6-sol
0.560
0.332
0.250
0.354
0.390
0.572
0.295
0.380
0.375
0.128
0.326
Opus 5.5
0.508
0.436
0.521
0.284
0.380
0.598
0.343
0.397
0.330
0.228
0.330
PulseLM
0.319
0.199
0.188
0.195
0.405
0.242
0.297
0.354
0.181
0.326
0.165
Table 2: Zero-shot physiological recognition. Macro-F1 ( ↑ ) for rhythm, quality (Qual.), perfusion (Perf.), morphology (Morph.), and respiratory-rate class (RR). Bold marks the best in each column.
Model
MC-MED
MIMIC-III
VitalDB
HR ↓
RR ↓
SBP ↓
MAP ↓
Fact ↑
HR ↓
RR ↓
Fact ↑
HR ↓
RR ↓
Fact ↑
GPT-5.6-luna
10.06
3.95
21.73
15.12
0.499
9.42
5.15
0.456
8.49
2.99
0.486
GPT-6-luna
11.24
3.93
22.24
14.51
0.559
11.52
4.71
0.533
8.86
2.34
0.579
GPT-6-sol
4.58
3.85
21.56
14.48
0.659
5.21
5.10
0.660
5.51
2.26
0.619
Opus 5.5
4.67
3.32
19.80
14.26
0.696
4.37
3.99
0.647
4.88
2.44
0.642
PulseLM
29.20
4.21
20.44
15.47
0.356
29.44
4.59
0.440
37.12
7.03
0.481
Table 3: Zero-shot physiological estimation. Fact is the accuracy of stated facts across six shared attributes ( ↑ ). Numerical columns report MAE ( ↓ ): HR in bpm, RR in breaths/min, and SBP/MAP in mmHg. MAE uses valid numerical outputs; bold marks the best reported value in each column.
Table 5: Zero-shot event-level physiological recognition. Class-versus-control AUROC ( ↑ ) for the five highest-scoring MC-MED classes for PPG-LM and all available VitalDB classes. Bold marks the best score in each column.
Model
MC-MED
VitalDB
CCB ↑
AntiHTN ↑
AntiPL ↑
AntiDM ↑
HTN ↑
DM ↑
Anemia ↑
eGFR ↑
GPT-5.6-luna
0.468
0.509
0.477
0.520
0.500
0.465
0.487
0.500
GPT-6-luna
0.497
0.484
0.421
0.464
0.475
0.507
0.495
0.498
GPT-6-sol
0.500
0.500
0.505
0.507
0.523
0.502
0.523
0.500
Opus 5.5
0.500
0.443
0.468
0.507
0.519
0.512
0.484
0.500
PulseLM
0.500
0.500
0.500
0.500
0.500
0.500
0.500
0.500
Table 6: Zero-shot visit-level recognition. Balanced accuracy ( ↑ ) for home-medication use (MC-MED) and clinical status (VitalDB). PPG-LM uses eight-window fact embeddings with held-out threshold calibration. Bold marks the best score in each column.
Task
Foundation models
PPG-LM
PaPaGei-S
PulsePPG
AnyPPG
SIGMA-PPG
MOMENT-L
Chronos-2
Stage 1
Stage 2
MC-MED (ED) — AUROC ↑
Sex
0.692 ± 0.006
0.719 ± 0.007
0.783 ± 0.008
0.712 ± 0.008
0.724 ± 0.005
0.739 ± 0.008
0.813 ± 0.006
0.856 ± 0.005
ED admission
0.673 ± 0.011
0.681 ± 0.009
0.750 ± 0.013
0.679 ± 0.011
0.700 ± 0.011
0.728 ± 0.013
0.765 ± 0.012
0.787 ± 0.010
Atrial fibrillation
0.742 ± 0.012
0.752 ± 0.013
0.845 ± 0.010
0.755 ± 0.013
0.768 ± 0.014
0.815 ± 0.012
0.857 ± 0.010
0.870 ± 0.011
MC-MED (ED) — R2↑
Table 7: Clinical prediction with linear probes. Three classification and three regression tasks per dataset (mean ± SD). Bold: best mean; † : window-level target.
Variant
Temporal acc. ↑
Event AUROC ↑
Segment R@1 ↑
w/o temporal loss
0.511
0.698
0.446
w/o signed offsets
0.923
0.617
0.445
w/o segment loss
0.922
0.697
0.202
w/o distillation
0.920
0.700
0.445
w/o seg. loss + distill.
0.919
0.699
0.015
PPG-LM
0.925
0.707
0.447
Table 8: Ablation studies.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Attribute
Evidence and description
Pulse rate
Median pulse-interval rate, expressed as a value and a category.
Rate trend
Instantaneous rate slope: falling, stable, or rising.
Rhythm
ECG R–R interval variability, expressed as a regularity category.
Perfusion
Pulsatile amplitude relative to baseline: weak, moderate, or strong.
Signal quality
NeuroKit2 quality score: clean, fair, or noisy but usable.
Augmentation
The pyPPG augmentation index: low, moderate, or high.
Appendix
Table 9: Segment-caption content. Perfusion and morphology are waveform surrogates rather than independent clinical diagnoses.
Component
Evidence and description
Demographics
Recorded age and sex.
Presentation
Arrival mode, Emergency Severity Index, and normalized chief complaints.
Triage measurements
Available heart rate, respiratory rate, oxygen saturation, blood pressure, and temperature.
Clinical history
Selected prior ICD-10 diagnoses, expressed as clinical descriptions.
Home medications
Recorded active home-medication classes.
Appendix
Table 10: Visit-caption content.
Component
Evidence and description
Event anchor
A recorded administration timestamp for an eligible medication class.
Pre-event context
The latest available PPG windows preceding the administration.
Post-event context
PPG windows sampled evenly across the period following administration.
Caption target
The medication class, with individual drug names, doses, and exact latencies omitted.
Control view
The same sampling rule around a within-visit anchor free of nearby eligible administrations.
Appendix
Table 11: Event-caption construction.
Operator
Evidence and description
Occurrence
Whether tachycardia or irregularity appears among known observations.
Proportion
The affected share of known windows, described as brief, part, or most.
Contiguous duration
Consecutive affected windows, grouped into short, intermediate, or long runs.
Rate ordering
Consecutive groups of observed rates, describing rising, falling, stable, rise-then-fall, or fall-then-rise patterns.
Appendix
Table 12: Temporal-caption content.
Split
Patients
Visits
30-second windows
Training
29,177
42,856
8,842,271
Validation
3,641
5,421
1,183,278
Test
3,612
5,196
1,099,395
Total
36,430
53,473
11,124,944
Appendix
Table 13: Processed MC-MED splits. Patients are disjoint across splits; window counts refer to the retained corpus.
Task
Reference and clinical interpretation
Rhythm
Regular, slightly irregular, or markedly irregular according to synchronized ECG R–R interval CV ( <0.05 , [0.05,0.12) , or ≥0.12 ). Measures beat-timing regularity, without assigning a specific arrhythmia.
Quality (Qual.)
Clean, fair, or noisy according to the annotation quality score and usability gate. Describes waveform reliability and artifact burden.
Perfusion (Perf.)
Weak, moderate, or strong pulsatile amplitude relative to the signal baseline, using the waveform-derived AC/DC index. A peripheral perfusion surrogate, not a direct blood-flow measurement.
Morphology (Morph.)
Low, moderate, or high augmentation from the pyPPG augmentation index. Describes pulse-wave shape; it does not establish a vascular disease diagnosis.
HR class
Five pulse-rate bins: <50 , 50–59, 60–100, 101–120, and >120 bpm, spanning markedly slow to fast rates. These are benchmark annotation categories.
HR estimation
Heart rate in beats/min, evaluated against synchronized ECG. This independent reference differs from the PPG-derived pulse rate used to construct training captions.
Appendix
Table 14: Physiological task definitions. Tasks in Tables 2 and 3 , including the HR-class target used in few-shot evaluation. CV denotes coefficient of variation.
Task
Data
Recorded target and clinical meaning
Administration events: medication class versus control
Rate control
M
Administration of a rate-controlling medication, used to slow heart rate.
Vasopressor
M, V
Administration of an agent used to increase arterial blood pressure.
Vasodilator
M, V
Administration of a medication that dilates blood vessels.
IV fluid
M
Intravenous fluid administration, representing fluid delivery to the circulation.
Diuretic
M
Administration of medication that promotes renal salt and water excretion.
Appendix
Table 15: Medication and visit-fact definitions. Dataset abbreviations refer to the tasks displayed in Tables 5 and 6 .
Task
Data
Definition and clinical interpretation
Shared demographic target
Sex
M, I, V
Recorded patient sex.
MC-MED: ED disposition and prior clinical history
ED admission
M
Hospital admission following the ED encounter, versus no admission.
Atrial fibrillation
M
Prior atrial fibrillation: disorganized atrial electrical activity associated with an irregular rhythm. The label records history, not necessarily AF in the observed window.
Table 16: Clinical classification targets. These rows cover all 15 dataset–task pairs in Table 22 ; sex occurs in all three datasets. Positive disease labels refer to recorded conditions.
Task
Data
Measurement and clinical interpretation
Demographics and body size
Age
M, I, V
Patient age in years. MIMIC-III values above 120 are treated as missing because of de-identification shifts.
BMI
V
Body mass index: weight divided by squared height (kg/m 2 ), a measure of body size relative to height.
Kidney function, hematology, and blood chemistry
eGFR
M
Estimated glomerular filtration rate (mL/min/1.73 m 2 ), reflecting kidney filtration function. This target is continuous.
Hemoglobin (Hb)
M, I
Blood concentration of the oxygen-carrying protein in red blood cells.
Appendix
Table 17: Clinical regression targets. These rows cover all 17 dataset–task pairs in Table 22 . Shared measurements are defined once; sampling times and target levels are specified in the accompanying text.
Rule-based evaluation
LLM-judge evaluation
Model / caption source
Fact acc. ↑
Coverage ↑
Factuality ↑
Unsupported ↓
[0pt][0pt] MC-MED (ED)
GPT-5.6-luna
0.499
0.99
2.63
6.54
PPG-LM : direct generation
0.923
1.00
3.64
2.94
PPG-LM : nearest-caption retrieval
0.877
1.00
3.83
1.51
[0pt][0pt] MIMIC-III (ICU)
Appendix
Table 19: Caption factuality by output method. Direct generation produces a new caption; nearest-caption retrieval selects an existing caption from the fixed bank. Fact accuracy covers six common fields; coverage includes all gradable fields. Judge factuality is on a 1–5 scale; unsupported claims are counted per caption. VitalDB uses the 40 Hz view. Bold marks the best value within each dataset.
(a) Triage-context recognition
Model
ESI ↑
Arrival ↑
HR ↑
SpO 2 ↑
RR ↑
Temp. ↑
BP ↑
CC ↑
GPT-5.6-luna
0.501
0.445
0.463
0.497
0.506
0.500
0.269
0.134
GPT-6-luna
0.500
0.498
0.486
0.500
0.500
0.500
0.238
0.128
GPT-6-sol
0.536
0.509
0.736
0.499
0.506
0.512
0.265
0.126
Opus 5.5
0.501
0.504
0.528
0.500
0.500
0.500
0.259
0.100
PulseLM
0.483
0.500
0.656
0.499
0.500
0.498
0.250
0.113
Appendix
Table 21: Additional visit-level clinical-fact recognition. Balanced accuracy ( ↑ ) for (a) triage context and (b) medical history on 400 MC-MED visits using eight windows each. PPG-LM read-outs use thresholds calibrated on held-out MC-MED validation visits. Bold marks column maxima.
Foundation models
PPG-LM
Task
PaPaGei-S
PulsePPG
AnyPPG
SIGMA-PPG
MOMENT-L
Chronos-2
Stage 1
Stage 2
MC-MED (ED) — AUROC ↑
Sex
0.692 ± 0.006
0.719 ± 0.007
0.783 ± 0.008
0.712 ± 0.008
0.724 ± 0.005
0.739 ± 0.008
0.813 ± 0.006
0.856 ± 0.005
ED admission
0.673 ± 0.011
0.681 ± 0.009
0.750 ± 0.013
0.679 ± 0.011
0.700 ± 0.011
0.728 ± 0.013
0.765 ± 0.012
0.787 ± 0.010
Atrial fibrillation
0.742 ± 0.012
0.752 ± 0.013
0.845 ± 0.010
0.755 ± 0.013
0.768 ± 0.014
0.815 ± 0.012
0.857 ± 0.010
0.870 ± 0.011
Heart failure
0.708 ± 0.018
0.723 ± 0.015
0.825 ± 0.013
0.711 ± 0.017
0.731 ± 0.015
0.781 ± 0.012
0.836 ± 0.010
0.860 ± 0.009
Appendix
Table 22: Full clinical prediction results. Test AUROC and R2 (mean ± SD; ↑ ) for all 32 tasks. All models use the same four windows for visit-level targets. Bold: best mean; † : window-level target.
Task
k=5
k=10
k=20
k=50
[0pt][0pt] MC-MED (ED) — Macro-F1 ↑
Rhythm
0.488 ± 0.051
0.537 ± 0.019
0.567 ± 0.037
0.597 ± 0.022
HR class
0.755 ± 0.023
0.831 ± 0.025
0.898 ± 0.018
0.921 ± 0.003
Sex
0.568 ± 0.042
0.645 ± 0.054
0.696 ± 0.041
0.710 ± 0.026
ED admission
0.563 ± 0.014
0.558 ± 0.026
0.583 ± 0.018
0.617 ± 0.008
[0pt][0pt] MIMIC-III (ICU) — Macro-F1 ↑
Appendix
Table 23: Few-shot adaptation of PPG-LM . Macro-F1 (mean ± SD; ↑ ) over five support draws. VitalDB uses the 40 Hz waveform view. k : examples per class. Bold: best mean per row.
Variant
Rhythm ↑
Quality ↑
R@1 ↑
[0pt][0pt] Stage 1
CLIP
0.650
0.505
0.790
SigLIP
0.620
0.400
0.815
CoCa
0.640
0.614
0.783
[0pt][0pt] Stage 2
CLIP
0.534
0.580
0.797
Appendix
Table 24: Objective ablations on MC-MED. Rhythm/quality: macro-F1; retrieval: signal-to-text R@1 with 2,000 candidates. Stage-2 CLIP/CoCa report per-metric bests across configurations.
MC-MED
MIMIC-III
VitalDB
Prompt bank
Rhythm
HR class
Quality
Rhythm
HR class
Quality
Rhythm
HR class
Quality
Template
0.651
0.355
0.673
0.670
0.337
0.786
0.389
0.327
0.583
Combined
0.629
0.393
0.592
0.642
0.350
0.677
0.384
0.282
0.515
Appendix
Table 25: Prompt sensitivity of PPG-LM . Template uses caption-style class descriptions; Combined adds reworded variants. Macro-F1 ( ↑ ); bold marks the best score per column.
Target
Template
Segment rate
The pulse rate is about {val} beats per minute, {level}.
Segment rhythm
The cardiac rhythm is {level}.
Segment quality
The PPG trace is {level}.
Poor-quality segment
This segment’s pulse waveform is too noisy to characterize reliably.
Medication event
During this period the patient was given {desc}.
Control interval
This period of monitoring contains no medication administration.
Appendix
Table 26: Representative training templates. Strings are reproduced verbatim; braced fields are filled from measurements or records.
Concept
Prompt
Pulse rate
Marked bradycardia
The measured pulse rate is markedly bradycardic.
Bradycardia
The measured pulse rate is bradycardic.
Normal rate
The measured pulse rate is within the normal range.
Mild tachycardia
The measured pulse rate is mildly tachycardic.
Tachycardia
The measured pulse rate is tachycardic.
Appendix
Table 27: Representative zero-shot prompts. Verbatim examples from the segment template bank and the visit and event class descriptions, grouped by task.
Photoplethysmography (PPG) plays a central role in wearable health monitoring and clinical decision support. Yet existing approaches to universal PPG representation learning largely focus on signal-level objectives and often overlook patient-level health context, which limits generalization to complex clinical tasks and heterogeneous cohorts. To address this gap, we construct a large-scale paired PPG-EHR multimodal dataset by distilling fragmented medical histories and clinical records into cohesive, patient-level electronic health records (EHR). Building on this resource, we propose Clinical Anchored Pretraining for PPG (CAP). During pretraining, CAP performs cross-modal contrastive alignment that anchors PPG representations to patient-level clinical semantics, guiding the encoder beyond waveform fitting toward modeling consistency in a patient's overall physiological state. During downstream adaptation, the pretrained PPG encoder provides clinically grounded representations that strengthen inductive bias and improve robustness and transferability. Experiments demonstrate that CAP consistently outperforms strong baselines on four diverse downstream tasks. CAP achieves a particularly large gain on respiratory rate prediction (up to +87.6% relative improvement over the state-of-the-art baseline) and delivers an average relative +26.7% across all tasks. We further enhance the interpretability of our approach through comprehensive analyses, including ablations and multiple complementary visualizations of the learned representations. The code for our experiments is available at: https://github.com/gody123gody/CAP .
Chenyang He, Xinyi Shao, Shun Huang +4
Nanjing University of Aeronautics and Astronautics Nanjing, Jiangsu, China · Peking University Beijing, Beijing, China · Independent Researcher +1
Photoplethysmography (PPG), a non-invasive measure of changes in blood volume, is widely used in both wearable devices and clinical settings. Recent PPG foundation models either use open-source ICU datasets with pretraining paradigms that require curated data and thus complicate generalization to field-like data, or use closed-source field-like PPG data. In contrast, we propose a PPG foundation model that does not require high-quality or field-like pretraining data, and instead leverages accompanying electrocardiogram and respiratory signals in ICU datasets to select contrastive samples during pretraining. Our approach allows the model to retain and learn from noisy PPG segments, improving robustness at inference. Our model, pretrained on 3x fewer subjects than existing state-of-the-art approaches, achieves performance improvements on 14 out of 15 diverse downstream tasks, including field-like daily activity and heart rate prediction. Our results demonstrate that multimodal supervision can integrate complementary physiological information to improve the robustness of PPG foundation models and enhance their generalization to consumer-grade data.
Eloy Geenjaar, Vince Calhoun, Scott Daly +4
Department of Electrical and Computer Engineering, Georgia Institute of Technology, Atlanta, USA · Tri-Institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS), Georgia State University, Georgia Institute of Technology, Emory University, Atlanta, USA · Dolby Laboratories, San Francisco, USA
Electrocardiography (ECG) is the clinical standard for cardiac assessment but requires dedicated hardware that does not scale to daily-life monitoring. Photoplethysmography (PPG) is ubiquitous in wearables but lacks ECG-specific diagnostic morphology and is corrupted by motion and sensor noise. PPG-to-ECG generation aims to bridge this gap by recovering electrical morphology and timing from peripheral pulse signals. However, existing methods largely rely on statistical alignment and data-driven generation. They fail to explicitly structure the latent space around physiology-aware electro-hemodynamic factors and lack constraints from forward physiological dynamics. To address these challenges, we propose PG-LRF, a physiology-guided latent rectified flow framework. PG-LRF introduces an electro-hemodynamic simulator that co-models ECG and PPG through shared cardiac phase dynamics. Guided by this simulator, a Physiology-Aware AutoEncoder learns a structured electro-hemodynamic latent space. Then we integrate this simulator guidance into a PPG-conditioned latent rectified flow, enforcing ECG-side morphology consistency and ECG-to-PPG forward hemodynamic consistency during generative transport. Experiments on the large-scale MC-MED dataset demonstrate that PG-LRF significantly improves PPG-to-ECG generation and downstream cardiovascular disease classification, proving its ability to generate ECGs that are both signal-faithful and physiologically plausible under the ECG-to-PPG hemodynamic pathway
Xiaoda Wang, Minxiao Wang, Kaiqiao Han +10
Department of Computer Science, Emory University · Department of Computer Science, University of California, Los Angeles · Nell Hodgson Woodruff School of Nursing, Emory University +2