cs.AISep 27, 2026

PPG-LM: A Photoplethysmography-Language Model with Multi-Level Clinical Alignment

Authors: Xiaoda Wang, Minxiao Wang, Maxwell A Xu, Patrick Langer, Kaiqiao Han, Defu Cao, Xiao Luo, Yuzhe Yang, +5 more

Organizations: Emory University · University of California, Los Angeles · Google · Stanford University · University of Southern California · University of Wisconsin–Madison

Abstract

Photoplethysmography (PPG) is widely recorded by clinical monitors and consumer wearables, providing a scalable source of continuous physiological information. These recordings offer an opportunity for physiological assessment at scale, but realizing this potential requires models to learn from both signal-derived physiological supervision and broader clinical context captured in electronic health records (EHRs). This involves aligning information spanning local observations, care events, and entire visits with PPG representations at corresponding temporal scales. However, existing PPG foundation models primarily rely on task-specific prediction heads, while the medical knowledge of large language models does not necessarily translate into waveform understanding. To bridge this gap, we introduce PPG-LM, the first PPG-language model family to learn physiological representations from both signal-derived supervision and broader clinical context captured in EHRs. To construct clinically grounded captions, we develop an automatic captioning pipeline that generates segment-, event-, and visit-level descriptions from signal measurements and structured EHR records. We then learn from these pairs through a two-stage framework that first establishes segment-language correspondence through contrastive learning and waveform-conditioned captioning, then extends alignment to events and visits through time-aware aggregation and temporal statement matching. Pretrained on approximately 73k hours of PPG, PPG-LM supports language-based recognition, cross-modal retrieval, and segment captioning. Experiments on MC-MED, MIMIC-III, and VitalDB show improved retrieval and caption factuality over language-model baselines and gains over PPG and time-series foundation models on multiple clinical prediction tasks.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CAP: Towards PPG Universal Representation Learning with Patient-level Supervision

    Jun 13, 2026Chenyang He, Xinyi Shao, Shun Huang +4Remote PhotoplethysmographyMedical World Model

  2. A robust PPG foundation model using multimodal physiological supervision

    Jun 5, 2026Eloy Geenjaar, Vince Calhoun, Scott Daly +4Remote PhotoplethysmographyHeart Rate

  3. PG-LRF: Physiology-Guided Latent Rectified Flow for Electro-Hemodynamic PPG-to-ECG Generation

    May 9, 2026Xiaoda Wang, Minxiao Wang, Kaiqiao Han +10CardiacPhysiological Data