PhysioTRACE: Provenance-Aware Stress Tests for Physiological Foundation Models
Organizations: Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, UAE · McGill University, Montreal, Canada · Carnegie Mellon University, Pittsburgh, USA
Abstract
Physiological foundation models encode how a signal was recorded alongside the physiology it reflects. When recording conditions are associated with diagnosis, this acquisition provenance can become a shortcut, yet the usual evidence, shifted transfer and provenance decodability, does not show whether a predictor uses it. We introduce PhysioTRACE, a four-axis behavioral audit for frozen encoders that separates what a probe can decode from what a fixed task head relies on. Recover scores how decodable provenance is; Stress reverses only the provenance-target association on the same held-out records; Intervene removes a train-localized provenance component; and Verify certifies that removal only if it beats matched random projections within a declared utility margin. Each audit thus ends in one of three verdicts: no reliance, or reliance with the remedy certified or refused. Across EEG and ECG, five training objectives, and five frozen foundation models, the relation between Recover's calibrated score and out-of-distribution utility changes sign between datasets, so neither can stand in for a reliance test. On paired EEG views where the shortcut is known by construction, the audit detects it (the exposed head loses about 0.2 AUROC when the association is reversed, while a control head is unaffected) and certifies removal of a rank-two component that restores control-level behavior without measurable utility loss, for both encoder objectives tested. On real ECG device metadata it returns all three verdicts: it certifies a remedy that removes 91% of one model's excess vulnerability, finds no reliance where device and diagnosis are barely associated, and refuses the remedy for a second model whose localized direction also carries task signal. Robustness to how inputs were recorded therefore needs a behavioral test, and PhysioTRACE provides one that can pass, fail, or refuse a remedy.
Figures & tables
| Paired EEG (NMT) | Observational ECG (PTB-XL devices) | ||||
| Task-only | Task + recon. | ECGFounder | ECGFounder, weak pair | CLEF-Small | |
| Excess vulnerability | 0.197 | 0.225 | 0.023 | 0.00007 | 0.060 |
| after suppression | 0.001 | 0.001 | 0.002 | – | 0.022 |
| Beats all random controls | yes | yes | yes | – | yes |
| Utility change | 0.000 | 0.001 | 0.004 | – | 0.039 |
| 95% lower bound | 0.004 | 0.001 | 0.006 | – | 0.044 |
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Removal operator | How the removed component is chosen | How removal is evaluated |
| INLP | Iterated null-space projection | Successive linear classifiers for the property | Property decodability after removal; downstream bias or task metrics |
| Amnesic Probing | INLP projection | Linear classifiers for the property | Change in the model’s own task behavior, against random directions of equal dimension |
| LEACE | Closed-form least-squares erasure | Cross-covariance of embedding and property | Guarantee that no linear classifier beats a constant predictor; downstream effects |
| PhysioTRACE | Whitened rank- projection (Eq. 4 ) | Paired within-record view contrasts, or a target-conditional probe fitted on training data | Added: fixed-marginal association stress on the same held-out records (Eq. 2 ); excess vulnerability against a control head fitted without the association (Eq. 3 ); paired same-record contrast (Eq. 7 ); fixed task heads; utility check against a declared margin. Reused: rank- and energy-matched random controls |
| Setting | Model / objective | OOD AUROC | Worst AUROC | LPI [95% interval] |
| ERP | Task-only | |||
| Augmentation-only | ||||
| Masked reconstruction | ||||
| CBraMod, frozen | ||||
| PTB | Task-only | |||
| Masked reconstruction |
| Benchmark | Train units/clusters | Validation units/clusters | Test units/clusters |
| ERP/P300 | |||
| PTB-XL |
| Model | ERP LPI [95% interval] | PTB LPI [95% interval] |
| Task-only | ||
| Augmentation-only | ||
| Invariance-first | ||
| Masked reconstruction | ||
| Supervised contrastive | ||
| LaBraM, frozen | n/a |
| Model / objective | LPI | Naive 95% interval | Reported 95% interval | Ratio | |
| ERP | Task-only | 0.485 | [0.474, 0.496] | [0.215, 0.702] | 22.8 |
| Augmentation-only | 0.503 | [0.491, 0.514] | [0.182, 0.754] | 24.8 | |
| Invariance-first | 0.237 | [0.229, 0.245] | [0.067, 0.367] | 19.0 | |
| Masked reconstruction | 0.917 | [0.908, 0.926] | [0.779, 1.028] | 14.3 | |
| Supervised contrastive | 0.361 | [0.352, 0.369] | [0.147, 0.548] | 22.9 | |
| LaBraM, frozen | 0.614 | [0.598, 0.630] | [0.380, 0.885] | 15.5 |
| ERP/P300 | PTB-XL | |||||
| Model | Acc. | BA | Macro-F1 | Acc. | BA | Macro-F1 |
| Task-only | 0.6819 | 0.5812 | 0.5802 | 0.4153 | 0.4154 | 0.4137 |
| Augmentation-only | 0.6924 | 0.5904 | 0.5889 | 0.4109 | 0.4110 | 0.4095 |
| Invariance-first | 0.5975 | 0.4872 | 0.4880 | 0.4140 | 0.4140 | 0.4126 |
| Masked reconstruction | 0.8017 | 0.7203 | 0.7193 | 0.3568 | 0.3567 | 0.3512 |
| Supervised contrastive | 0.6446 | 0.5332 | 0.5355 | 0.4020 | 0.4021 | 0.3990 |
| Objective | Seed | Probe, original | Probe, refitted | ||
| Task-only | 7 | ||||
| 13 | |||||
| 23 | |||||
| 31 | |||||
| 47 | |||||
| Task + reconstruction | 7 |
| Encoder seed | Task-only LPI | Reconstruction LPI | Difference | interval |
| 13 | ||||
| 23 | ||||
| 31 | ||||
| 47 | ||||
| Four-seed mean |
| Head | Representation | Aligned | Independent | Reversed | Flip rate |
| Control | full | ||||
| Control | targeted rank-2 removed | ||||
| Exposed | full | ||||
| Exposed | random rank-2 removed | ||||
| Exposed | targeted rank-2 removed |
| Head | Representation | Aligned | Independent | Reversed | Flip rate |
| Control | full | ||||
| Control | targeted rank-2 removed | ||||
| Exposed | full | ||||
| Exposed | random rank-2 removed | ||||
| Exposed | targeted rank-2 removed |
| Model | Head | Representation | Aligned ( ) | Indep. ( ) | Reversed ( ) | ||
| ECGFounder | Exposed | full | |||||
| Exposed | projected | ||||||
| Control | full | ||||||
| Control | projected | ||||||
| CLEF-Small | Exposed | full | |||||
| Exposed | projected |
| Gate | Criterion | Estimate | Result |
| Primary excess gap | and 95% lower bound | pass | |
| Control reversed advantage | and 95% lower bound | pass | |
| Exposed aligned noninferiority | 95% lower bound | pass | |
| Conditional device recoverability | 95% lower bound | pass | |
| Embedding probe above coarse metadata | paired 95% lower bound | pass | |
| Targeted projection reduces | 95% lower bound | pass |
| Frozen encoder | Same-reference BA | Shifted BA | Worst-reference BA | Shift loss | Reference-probe BA |
| BIOT | |||||
| CBraMod | |||||
| LaBraM |
| Sensitivity | Result (95% interval) | Interpretation |
| Full-cohort metadata comparator | Metadata age/sex/year | Calendar and workflow are material confounds; the primary embedding result is not a hardware effect beyond these variables. |
| Calendar-matched cohort (1990–1999) | ECGFounder ; paired excess over metadata | A coarse linear calendar comparator does not reproduce the entire embedding effect on this shared support; nonlinear era/workflow confounding remains possible. |
| Human-validated ECGs | The direction is positive but below the primary 0.01 meaningful-effect gate; sparse cells make this a sensitivity check. | |
| Weak-association site-2 device pair | The interval lies within under fold-matched stress, consistent with association specificity rather than a generic reweighting effect. |
| Encoder | Variant | OOD AUROC | Worst OOD AUROC | Domain probe |
| EEGNetSmall | Masked reconstruction | |||
| EEGNetSmall | Augmentation only | |||
| EEGNetSmall | Invariance first | |||
| ShallowConvNet | Masked reconstruction | |||
| ShallowConvNet | Augmentation only | |||
| ShallowConvNet | Invariance first |
| Paired EEG (NMT) | Observational ECG (PTB-XL) | |||
| Task-only | Task + recon. | ECGFounder | CLEF-Small | |
| Recover | 1.18 [1.09, 1.25] | 1.27 [1.16, 1.37] | 0.959 [0.953, 0.964] | 0.897 [0.887, 0.906] |
| Excess vulnerability | 0.197 [0.173, 0.221] | 0.225 [0.193, 0.256] | 0.023 [0.021, 0.026] | 0.060 [0.055, 0.065] |
| after suppression | 0.001 [ 0.003, 0.002] | 0.001 [ 0.002, 0.004] | 0.002 [0.001, 0.003] | 0.022 [0.020, 0.025] |
| Beats random controls | 10/10 per seed | 10/10 per seed | 100/100 | 100/100 |
| Utility change | 0.000 [ 0.004, 0.004] | 0.001 [ 0.001, 0.005] | 0.004 [ 0.006, 0.001] | 0.039 [ 0.044, 0.033] |
| Factor role | Examples | Desired representation behavior | Evaluation implication |
| Near nuisance | Valid re-referencing, geometry-aware channel subsets, benign resampling, or bounded acquisition changes known to preserve physiology and target semantics | Preserve task-relevant physiology while limiting unstable acquisition identity | Use matched views when available; require shifted and worst-group utility, fixed-head verification, and a utility margin before suppression claims. |
| Label-informative | An acquisition or protocol variable that is part of the intended measurement and has domain-supported, stable deployment meaning | Model the factor explicitly rather than imposing invariance | State the operational rationale and audit whether its predictive role remains stable across supported deployment groups. |
| Confounded | Site, device, protocol, or workflow differences that also track population, label prevalence, or clinical practice without matched counterfactual views | Separate accessibility from decision reliance and avoid biological or hardware-causal interpretation | Use identity-disjoint splits, composition reporting, target-conditional probes, fixed-marginal stress, and control heads fitted without the association. |
| Shift family | Controlled stress test | Natural held-out group | Instantiation and interpretation here |
| Reference and montage | Valid re-referencing, reference mixing, channel dropping, subset sampling, or geometry-aware remapping | Held-out reference family, montage family, channel set, or reference pipeline | NMT is the controlled anchor: the task is fixed and the same records are evaluated under AVG, CZREF, and LE reference views. |
| Device, site, and protocol | Device-aware perturbations, acquisition-pipeline variants, or protocol-specific preprocessing changes when physically justified | Held-out amplifier, device model, collection site, task interface, or clinical workflow | ERP/P300 tests real-domain EEG transfer across interface and population differences; PTB-XL tests clinical ECG site transfer. Both require confound-aware interpretation. |
| Sampling, filters, and timing | Resampling with anti-aliasing, bounded filter-chain perturbations, mild line-noise variation, or clock-drift-like jitter | Held-out sampling rate, filter chain, artifact-removal policy, or segmentation convention | Included in the provenance schema and leakage audit. The present experiments do not claim causal isolation of filter or timing effects. |
| Quality and environment | Realistic impedance-like noise, dropped-channel patterns, sensor-contact variation, or bounded motion/artifact perturbations | Held-out quality stratum, recording environment, wearable condition, or artifact regime | Treated as a reporting requirement here; future datasets can instantiate this family without changing the protocol metrics. |
| Category | Fields to record or standardize |
| Device/amplifier | Manufacturer, model, amplifier, analog-to-digital converter (ADC) characteristics, firmware or acquisition software when available |
| Reference or lead definition | Online reference, re-reference procedure, lead configuration, polarity conventions |
| Montage or sensor geometry | Channel names, electrode positions, lead set, missing-channel policy, remapping or interpolation procedure |
| Sampling and timing | Sampling rate, anti-alias settings, clock drift information, resampling procedure |
| Filters and preprocessing | Hardware and software filter chain, notch settings, artifact removal, independent component analysis (ICA) or artifact subspace reconstruction (ASR) flags, normalization, segmentation policy |
| Quality and impedance | Per-channel impedance or quality indicators when available, dropped channels, sensor-contact metadata |
| Artifact component | Purpose |
| README.md | High-level artifact overview, repository layout, quickstart, data requirements, and expected output files. |
| ARTIFACT.md | Exact lightweight verification commands, data setup instructions, full-rerun entry points, and result-source map. |
| PROTOCOL_CHECKLIST.md | Compact checklist for applying provenance-aware evaluation: acquisition groups, split unit, leakage controls, shift interpretation, required metrics, and saved outputs. |
| pyproject.toml , requirements.txt , LICENSE | Installable package metadata, Python dependency list, and software license. |
| sharable_modules/ | Reusable dataset registry, split preparation, preprocessing, model primitives, training helpers, frozen-probe utilities, transfer-matrix construction, pooled-target metrics, and common reporting semantics. |
| benchmarks/ | Benchmark-specific modules layered above the shared utilities, including NMT reference-shift reporting, ERP/P300 domain-transfer reporting, PTB-XL site-shift reporting, external LaBraM/ECGFounder frozen-encoder utilities, and shared report-export mechanics. |
| Signal | Model | Pretraining corpus coverage | Multi-ch. | Multi-dev. |
| EEG | REVE ( Ouahidi et al., 2025 ) | 92 public EEG datasets; approximately 25,000 subjects; designed for adaptation across heterogeneous EEG setups. | Yes | Yes |
| EEG | EEGPT ( Wang et al., 2024 ) | PhysioMI/EEGMMIDB, HGD, TUH EEG, SEED, and M3CV; broad EEG corpus spanning motor, clinical, emotion, and multi-session settings. | Yes | Yes |
| EEG, ECG | BIOT ( Yang et al., 2023 ) | SHHS, CHB-MIT, IIIC Seizure, TUAB, TUEV, and HAR-style biosignal or wearable sources; designed for cross-data learning in the wild. | Yes | Yes |
| EEG | BENDR ( Kostas et al., 2021 ) | TUEG/TUH EEG Corpus; large-scale EEG pretraining with a wav2vec-style contrastive objective. | No | Yes |
| EEG | CBraMod ( Wang et al., 2025 ) | TUEG/TUH EEG Corpus; criss-cross EEG foundation-model pretraining. | No | Yes |
| EEG | LaBraM ( Jiang et al., 2024 ) | Large multi-dataset EEG pretraining corpus; discrete neural-token modeling for generic EEG representations. | Yes | Yes |
| Signal | Model | Preprocessing signature | Type | Loss or objective |
| EEG | REVE ( Ouahidi et al., 2025 ) | four-dimensional spatiotemporal positional encoding using electrode coordinates and time index; patching without requiring a fixed montage. | R | Masked autoencoding: reconstruct masked raw patches with an term plus a weighted global-token term. |
| EEG | EEGPT ( Wang et al., 2024 ) | Patch multichannel EEG; channel-identity embeddings and temporal rotary position encoding; masks over time and channels. | H | Alignment to a momentum target plus masked reconstruction with MSE-style losses on normalized targets. |
| EEG, ECG | BIOT ( Yang et al., 2023 ) | Resampling, per-channel amplitude normalization, fixed-length tokenization, FFT-energy features, and channel plus relative-position embeddings. | C | Contrastive alignment of sample embeddings under token and channel dropout, optimized with a similarity-matrix cross-entropy loss. |
| EEG | BENDR ( Kostas et al., 2021 ) | Scale and shift to a common range, resample to 256 Hz, map to 19 10/20 channels with missing channels zero-filled, and use 60 s segments. | C | wav2vec2-style contrastive loss on masked latents with in-sequence negatives plus an activation penalty. |
| EEG | CBraMod ( Wang et al., 2025 ) | Convolutional feature encoder plus transformer; time- and frequency-aware patch features with patch masking. | R | Masked patch reconstruction with MSE-style losses in time/frequency feature space. |
| EEG | LaBraM ( Jiang et al., 2024 ) | Vector-quantized neural tokenizer via spectrum prediction; per-sample normalization; patches over channels and time. | R | Discrete masked-token prediction over learned neural codes. |
| Evaluation | Groups | Metric | Seeds | Appendix role |
| NMT clinical EEG reference shift | AVG, CZREF, LE | balanced accuracy | 7, 13, 23, 42, 52 | controlled EEG reference-shift evaluation |
| ERP/P300 domain transfer | ALS P300, covert GeoSpell, overt P300 | AUROC | 7, 13, 23, 42, 52 | primary real-domain EEG transfer evaluation |
| ERP encoder check | ALS P300, covert GeoSpell, overt P300 | AUROC | 7, 13, 23, 42, 52 | architecture robustness check |
| NMT frozen EEG encoders (LaBraM, BIOT, CBraMod) | AVG, CZREF, LE | balanced accuracy | 7, 13, 23, 42, 52 | external frozen EEG foundation-model audits |
| ERP frozen EEG encoders (LaBraM, BIOT, CBraMod) | ALS P300, covert GeoSpell, overt P300 | AUROC | 7, 13, 23, 42, 52 | external frozen EEG foundation-model audits |
| PTB-XL cross-site transfer | site 0, site 1, site 2 | AUROC | 7, 13, 23, 42, 52 | cross-modality clinical site-shift evaluation |
| Evaluation | Contrast | Shifted/OOD | Worst | Gap | Probe |
| NMT reference | Contrastive canonical reconstruction + invariance | 0.054 0.019 | 0.071 0.030 | -0.046 0.009 | -0.221 0.086 |
| NMT reference | Supervised contrastive canonical reconstruction + invariance | 0.056 0.030 | 0.090 0.038 | -0.045 0.015 | -0.191 0.101 |
| ERP/P300 domains | Augmentation masked reconstruction | 0.197 0.028 | 0.241 0.046 | -0.063 0.038 | -0.110 0.035 |
| ERP/P300 domains | Invariance masked reconstruction | 0.172 0.021 | 0.222 0.047 | -0.064 0.038 | -0.218 0.017 |
| PTB-XL sites | Task only masked reconstruction | 0.188 0.050 | 0.217 0.048 | -0.010 0.006 | 0.038 0.067 |
| PTB-XL sites | Augmentation masked reconstruction | 0.188 0.049 | 0.216 0.050 | -0.011 0.008 | 0.010 0.065 |
| Variant | AVG target | CZREF target | LE target |
| AVG-only | 0.666 0.017 | 0.543 0.052 | 0.568 0.049 |
| Mixed refs | 0.682 0.008 | 0.682 0.004 | 0.678 0.009 |
| Contrastive only | 0.652 0.006 | 0.654 0.013 | 0.662 0.010 |
| Supervised contrastive | 0.632 0.017 | 0.635 0.020 | 0.638 0.024 |
| Canonical recon. + invariance | 0.639 0.041 | 0.636 0.050 | 0.634 0.048 |
| LaBraM frozen | 0.688 0.011 | 0.628 0.010 | 0.656 0.014 |
| Source | Target | Contrastive only | Supervised contrastive | Canonical recon. + invariance | LaBraM frozen |
| AVG | AVG | 0.635 0.022 | 0.631 0.016 | 0.614 0.040 | 0.647 0.042 |
| AVG | CZREF | 0.639 0.016 | 0.626 0.022 | 0.531 0.035 | 0.567 0.032 |
| AVG | LE | 0.639 0.018 | 0.631 0.030 | 0.581 0.054 | 0.596 0.042 |
| CZREF | AVG | 0.624 0.040 | 0.634 0.020 | 0.585 0.029 | 0.681 0.012 |
| CZREF | CZREF | 0.623 0.033 | 0.634 0.021 | 0.628 0.037 | 0.641 0.023 |
| CZREF | LE | 0.630 0.039 | 0.634 0.025 | 0.596 0.036 | 0.671 0.022 |
| Source | Target | Masked recon. | Augmentation only | Invariance first | LaBraM frozen |
| ALS P300 | ALS P300 | 0.632 0.029 | 0.804 0.016 | 0.764 0.019 | 0.536 0.024 |
| ALS P300 | Covert GeoSpell | 0.630 0.030 | 0.757 0.020 | 0.739 0.008 | 0.587 0.022 |
| ALS P300 | Overt P300 | 0.701 0.069 | 0.922 0.006 | 0.901 0.005 | 0.586 0.028 |
| Covert GeoSpell | ALS P300 | 0.534 0.070 | 0.798 0.016 | 0.760 0.018 | 0.498 0.020 |
| Covert GeoSpell | Covert GeoSpell | 0.647 0.040 | 0.757 0.019 | 0.740 0.008 | 0.613 0.009 |
| Covert GeoSpell | Overt P300 | 0.688 0.118 | 0.919 0.006 | 0.900 0.006 | 0.659 0.030 |
| Source | Target | Masked recon. | Task only | Augmentation only | ECGFounder frozen |
| Site 0 | Site 0 | 0.671 0.050 | 0.856 0.011 | 0.856 0.009 | 0.822 0.002 |
| Site 0 | Site 1 | 0.705 0.078 | 0.910 0.021 | 0.912 0.015 | 0.906 0.002 |
| Site 0 | Site 2 | 0.723 0.053 | 0.907 0.016 | 0.910 0.016 | 0.910 0.002 |
| Site 1 | Site 0 | 0.664 0.064 | 0.860 0.006 | 0.857 0.006 | 0.791 0.004 |
| Site 1 | Site 1 | 0.730 0.072 | 0.924 0.004 | 0.922 0.004 | 0.899 0.002 |
| Site 1 | Site 2 | 0.732 0.075 | 0.914 0.003 | 0.917 0.002 | 0.893 0.003 |
| Asset | Source / citation | License or terms | Use in this paper |
| NMT Scalp EEG Dataset | Khan et al. ( Khan et al., 2022 ) ; official NMT dataset/code page | Dataset: CC BY-SA 4.0; associated code repository: BSD 3-Clause | Used for the controlled EEG reference-shift evaluation across AVG, CZREF, and linked-ear reference views. |
| PTB-XL ECG Dataset | Wagner et al. ( Wagner et al., 2020a ) ; PhysioNet version 1.0.1 | CC BY 4.0 | Used for the ECG cross-site stress test with patient-level official folds and site-defined provenance groups. |
| Brain/Neural Computer Interaction (BNCI) 2014-008 ALS P300 Dataset | Riccio et al. ( Riccio et al., 2013 ) ; BNCI Horizon 2020 dataset 008-2014 | CC BY-NC-ND 4.0 | Used as the ALS P300 domain in the ERP/P300 real-domain transfer evaluation. |
| BNCI 2014-009 Covert and Overt ERP-based brain-computer interface (BCI) Dataset | Aricò et al. ( Aricò et al., 2014 ) ; Aloise et al. ( Aloise et al., 2012 ) ; BNCI Horizon 2020 dataset 009-2014 | CC BY-NC-ND 4.0 | Used for the covert GeoSpell and overt P300 domains in the ERP/P300 transfer evaluation. |
| LaBraM checkpoint and code | Jiang et al. ( Jiang et al., 2024 ) ; official LaBraM repository | MIT License | Used only as a frozen EEG encoder for the NMT, ERP/P300, and HMC evaluations; no LaBraM checkpoint is redistributed in our artifact. |
| ECGFounder checkpoint and code | Li et al. ( Li et al., 2025 ) ; official ECGFounder repository/model card | MIT License | Used only as a frozen ECG encoder for the PTB-XL audit; no ECGFounder checkpoint is redistributed in our artifact. |