Gaussian head avatars typically model intrinsic facial appearance as temporally static, omitting subtle cardiac-induced skin-color variation. We propose Heartian, a physiology-aware modulation framework that learns cardiac-cycle-dependent per-frame albedo modulation of facial skin-region Gaussians within a relightable head avatar to encode remote photoplethysmography (rPPG) signals. Using synchronized contact PPG supervision, Heartian models the prescribed cardiac waveform as the sum of two Gaussian functions and learns per-frame spatial residuals via a lightweight MLP. Across 152 stationary recordings from UBFC-rPPG, PURE, and MMPD, attribute-space recovery of the supplied signal achieves a pooled recording-level heart-rate MAE of 0.29 bpm and MAPE of 0.38%. The signals remain detectable after rendering by benchmark rPPG methods, with the best tested configuration - a motion-augmented TS-CAN decoder pretrained on UBFC-rPPG - recovering heart rate from the rendered MMPD avatars at 0.97 bpm MAE and 1.21% MAPE. Meanwhile, Heartian maintains reconstruction quality comparable to the baseline, with negligible average PSNR degradation of 0.005 dB. Overall, our work embeds recoverable rPPG signals as controllable material attributes to subject-specific Gaussian head avatars while retaining the reconstruction quality.
Figures & tables
Figure 1. Skin Gaussians Gs (red) overlaid on G (blue).
Figure 2. A typical PPG waveform with systolic (blue) and diastolic waves (red) in the time (a) and phase domains (b).
Test Set
UBFC-rPPG
PURE
MMPD
Method
Train Set
MAE ↓
MAPE ↓
MAE ↓
MAPE ↓
MAE ↓
MAPE ↓
POS
-
1.23 ± 0.77
1.01 ± 0.62
14.23 ± 6.78
29.34 ± 13.98
1.53 ± 0.63
2.66 ± 1.20
TS-CAN
UBFC-rPPG
-
-
10.01 ± 4.52
17.89 ± 8.68
2.59 ± 0.52
3.89 ± 0.94
PURE
9.22 ± 7.47
7.41 ± 5.56
-
-
2.51 ± 0.57
3.87 ± 1.03
TS-CAN (MA)
UBFC-rPPG
-
-
4.74 ± 4.22
5.23 ± 4.67
0.97 ± 0.27
1.21 ± 0.31
Table 1. Benchmark Evaluation Results. Performance of benchmark methods on our rendered Heartian videos, evaluated across datasets. Green highlights indicate metrics where embedded signals remain recoverable on par with the original benchmark results on the same selected subsets.
Baseline
Ours
Dataset
MAE
MAPE
SNR
MAE ↓
MAPE ↓
SNR ↑
UBFC-rPPG
59.00
53.81
-21.52
0.00
0.00
1.45
PURE
21.53
25.37
-7.13
0.17
0.27
3.47
MMPD
39.11
34.81
-11.47
0.33
0.42
4.04
Table 2. Attribute-level Evaluation. rPPG signal metrics extracted from the Gaussian albedo of baseline and Heartian.
POS
TS-CAN (MA)
FactorizePhys
MAE ↓
MAPE ↓
MAE ↓
MAPE ↓
MAE ↓
MAPE ↓
Skin tone
3
1.01 ± 0.82
1.80 ± 1.52
1.04 ± 0.44
1.35 ± 0.52
2.89 ± 1.48
4.66 ± 2.38
4
1.91 ± 1.44
2.33 ± 1.82
2.17 ± 1.11
2.35 ± 1.11
2.89 ± 1.88
3.61 ± 2.41
5
3.80 ± 2.34
7.44 ± 4.76
0.62 ± 0.34
0.87 ± 0.45
4.90 ± 2.89
9.50 ± 5.81
6
0.47 ± 0.21
0.53 ± 0.25
0.10 ± 0.10
0.12 ± 0.11
0.10 ± 0.07
0.12 ± 0.09
Table 3. Lights and skin tones comparison within MMPD ( Tang et al., 2023 ) , tested with models pretrained on UBFC-rPPG ( Bobbia et al., 2019 ) .
Test Set
UBFC-rPPG
PURE
MMPD
Method
Train Set
MAE
RMSE
MAPE
ρ
SNR
MAE
RMSE
MAPE
ρ
SNR
MAE
RMSE
MAPE
ρ
SNR
POS
-
1.23 ± 0.77
2.75 ± 2.43
1.01 ± 0.62
0.98 ± 0.05
0.26 ± 0.92
14.23 ± 6.78
25.75 ± 18.04
29.34 ± 13.98
0.48 ± 0.30
2.93 ± 0.56
1.53 ± 0.63
7.43 ± 5.53
2.66 ± 1.20
0.82 ± 0.04
4.20 ± 0.25
TS-CAN
UBFC-rPPG
-
-
-
-
-
10.01 ± 4.52
17.46 ± 12.54
17.89 ± 8.68
0.68 ± 0.25
-3.31 ± 0.66
2.59 ± 0.52
6.60 ± 4.60
3.89 ± 0.94
0.85 ± 0.04
-1.71 ± 0.17
PURE
9.22 ± 7.47
25.37 ± 24.62
7.41 ± 5.56
0.23 ± 0.34
-4.9 ± 1.14
-
-
-
-
-
2.51 ± 0.57
7.02 ± 4.85
3.87 ± 1.03
0.83 ± 0.04
-1.76 ± 0.16
TS-CAN (MA)
UBFC-rPPG
-
-
-
-
-
4.74 ± 4.22
14.18 ± 13.80
5.23 ± 4.67
0.82 ± 0.19
-0.28 ± 0.96
0.97 ± 0.27
3.35 ± 2.63
1.21 ± 0.31
0.96 ± 0.02
-0.40 ± 0.15
Table 1. Additional Benchmark Evaluation Results. Performance of benchmark methods on our rendered Heartian avatar videos, evaluated across the UBFC-rPPG ( Bobbia et al., 2019 ) , PURE ( Stricker et al., 2014 ) , and MMPD ( Tang et al., 2023 ) datasets. Green highlights indicate metrics where signals remain recoverable on par with the benchmark results on the same subsets.
Figure 1. Visualization of Recovered Heart Rate Accuracy. Scatter plots of rPPG estimates from our rendered Heartian videos versus ground truth using POS for UBFC-rPPG and MMPD, and motion-augmented TS-CAN for PURE and MMPD.
Test Set
UBFC-rPPG
PURE
MMPD
Method
Train Set
MAE
RMSE
MAPE
ρ
SNR
MAE
RMSE
MAPE
ρ
SNR
MAE
RMSE
MAPE
ρ
SNR
POS
-
+ 0.87
+ 1.15
+ 0.87
- 0.01
- 0.94
+ 0.52
+ 0.05
+ 0.62
- 0.01
- 0.22
+ 3.04
+ 5.45
+ 5.62
- 0.24
- 2.09
TS-CAN (MA)
UBFC-rPPG
-
-
-
-
-
0.00
0.00
0.00
0.00
+ 0.06
- 0.30
- 1.90
+ 0.30
+ 0.01
+ 0.04
FactorizePhys
UBFC-rPPG
-
-
-
-
-
- 0.17
- 0.30
- 0.38
+ 0.01
- 0.11
- 0.05
- 0.01
- 0.06
0.00
- 0.10
PURE
0.00
0.00
0.00
0.00
- 0.05
-
-
-
-
-
- 1.37
- 4.47
- 2.90
+ 0.13
- 0.07
Table 2. Background Compositing Ablation. Metric changes from background compositing over the default black background.
Figure 2. Relighting Ablation Visualization. From left to right: white, cool, warm, and real-world hospital lighting ( Source Link ); two subjects from PURE (left) and UBFC-rPPG (right). Solid color environment maps are synthesized with matched luminance.
Figure 3. Ablation on Spatial Residual of Our Modulation Components. Both ours (blue) and GT (orange) are filtered and normalized. Full modulation shows a better alignment of the prescribed signals, with 0.05 MACC improved in this case.
Remote photoplethysmography (rPPG) enables non-contact heart-rate estimation from facial videos, but its weak physiological signal is easily corrupted by motion, illumination changes, occlusion, skin-appearance variation, and device noise. Existing rPPG methods typically rely on a single model to directly predict heart rate or recover pulse waveforms, while different strong estimators may produce conflicting yet individually plausible candidates for the same video. To resolve these conflicts, we propose PhysAgent, an inference-time multi-agent candidate-verification framework. Unlike direct prediction approaches, PhysAgent neither trains a new base rPPG model nor asks Multimodal Large Language Models (MLLMs) to output heart rate directly. In contrast, it treats outputs from multiple base estimators as physiological hypotheses to be verified and uses a lightweight 4B MLLM, Qwen3-VL-4B, to drive multi-agent reasoning over video conditions, signal reliability, and candidate disagreement. A deterministic physiological verifier checks the fusion proposal, and a reproducible numerical fusion process produces the final heart rate. Experimental results on multiple public rPPG benchmarks show that PhysAgent improves fusion stability and reliability across different datasets and source-domain settings, while avoiding the irreproducibility and physiological inconsistency of direct MLLM prediction or unconstrained ensemble fusion. The code will be released soon.
Yehui Yang, Bo Zhao, Junzhe Cao +4
1Great Bay University · 2Shenzhen University of Advanced Technology · 3Macao Polytechnic University +1
Deep remote photoplethysmography (rPPG) attains sub-bpm heart-rate error on frontal, stationary faces yet degrades sharply under head pose: on MMPD, the state-of-the-art FactorizePhys backbone's MAE grows 1.60× from frontal (∣yaw∣<15∘) to large-yaw (∣yaw∣≥45∘) frames. We argue that pose is a \emph{coordinate-structural} nuisance rather than a data-augmentation problem: in image coordinates the same pixel maps to different anatomy at different poses, blocking three priors otherwise natural for rPPG, namely the dichromatic reflection model, pulse-phase invariance across skin regions, and the POS/CHROM chromaticity projection, each of which presumes a stable anatomy-to-pixel mapping. We introduce \textbf{CanonicalPhys}, which prepends a differentiable four-point homography that fixes four facial anchors at canonical positions; in this canonical frame the three priors become expressible as a per-pixel Lambertian weight, a cross-ROI temporal consistency loss, and knowledge distillation from windowed POS, none of which adds trainable parameters over the backbone. At an identical parameter count, CanonicalPhys reduces MMPD's frontal-to-large-yaw MAE degradation from 1.60× to 1.33× and flattens the mild-yaw bin from 1.32× to 1.07× (across CanonicalPhys variants), with matched cross-dataset MAE reductions of up to 32% on pose-rich targets. Code: https://github.com/infraface/CanonicalPhys
Hui Wei, Seyedata Jodeiri Seyedian, Xiaobai Li +1
Center for Machine Vision and Signal Analysis (CMVS), University of Oulu · ELLIS Institute Finland · Zhejiang University
Objective: Consumer face camera remote photoplethysmography (rPPG) enables passive cardiovascular monitoring, but whether single-cycle waveform morphology encoding arterial stiffness biomarkers is recoverable from this measurement has not been characterised. Methods: We evaluated 16 architectures spanning six families on 153 subjects across three datasets, introducing cross-subject Pearson r to distinguish subject-specific recovery from template collapse. Results: No architecture recovered subject-specific morphology (cross-subject r range 0.773--0.9999; ground-truth ceiling 0.601). Supervised Contrastive (SupCon) converged to log N = 4.844, constituting the strongest available empirical evidence that no discriminative morphological structure is extractable from single-cycle rPPG by the encoder families tested. The VAE decoder restores population-level harmonic content absent from the rPPG input (H2/H1: 0.310 output vs. 0.275 input), generalising zero-shot to UBFC (r = +0.708); a directional hallucination gap (p = 0.150) suggests partial signal reading. Anti-collapse objectives fail when input carries no discriminative structure. Significance: Consumer cameras cannot encode individual arterial morphology; cross-subject r is a necessary collapse diagnostic for waveform reconstruction benchmarks.