PHASE: A Physiology-Guided Hierarchical Foundation Model for Intracranial EEG
Organizations: UCLA Samueli School of Engineering · UCLA Mattel Children’s Hospital, David Geffen School of Medicine · Children’s Hospital of Michigan, Wayne State University School of Medicine · University of Pennsylvania
Abstract
Clinicians and neuroscientists have long analyzed intracranial electroencephalography (iEEG) through directly measurable physiological characteristics, which carry much of the information that downstream tasks depend on. Recent iEEG foundation models learn by reconstructing or predicting their inputs, which leaves the retention of these characteristics implicit. They are also evaluated mainly on cognitive decoding and a narrow clinical task, i.e., seizure detection. On a broad, clinically relevant benchmark such as Omni-iEEG, they remain below task-specific models when used frozen. We introduce PHASE, a physiology-guided foundation model that makes these characteristics explicit learning targets, pairing them with masked latent prediction in a temporal stage (PHASE-T) within each channel and a spatiotemporal stage (PHASE-ST) across synchronized channels. PHASE is pretrained on heterogeneous recordings from 222 participants at nine clinical sites. On all five Omni-iEEG clinical tasks, frozen PHASE-T outperforms every evaluated foundation model by up to 31%, and fine-tuned PHASE-T surpasses the task-specific models, setting a new state of the art. PHASE-T benefits from physiological supervision, outperforming variants trained with latent prediction alone or auxiliary waveform reconstruction on every task in matched ablations. PHASE-T generalizes to unseen institutions, outperforming the compared models with few or no local labels. PHASE-ST further improves seizure-onset-zone identification over PHASE-T and, when frozen, decodes sound volume and pitch on BrainTreebank better than published models. Beyond task performance, PHASE learns to encapsulate the physiological characteristics clinicians recognize, from seizure onset and its propagation to anatomical region identity, even though its pretraining contains no ictal recordings or anatomical labels.
Figures & tables
| Event | Window (60 s) | Anatomy | Pathol. brain region | ||||
| Model | HFO | Ictal | Sleep | Lobe-5 | Region-12 | Channel | Outcome |
| F1 | F1 | F1 | F1 | F1 | AUC | AUC | |
| PyHFO-Omni | 0.8061 | – | – | – | – | 0.7351 | 0.7438 |
| CLAP | – | 0.9245 | 0.7225 | 0.4750 | 0.3540 | 0.7684 | 0.6770 |
| TimeConv-CNN | – | 0.8533 | 0.7118 | 0.4788 | 0.3087 | 0.8061 | 0.7380 |
| SEEG-NET | – | 0.7526 | 0.6773 | 0.2520 | 0.1081 | 0.7850 | 0.5952 |
| Event | Window (60 s) | Anatomy | Pathol. brain region | ||||
|---|---|---|---|---|---|---|---|
| Pretraining variant | HFO | Ictal | Sleep | Lobe-5 | Region-12 | Channel | Outcome |
| F1 | F1 | F1 | F1 | F1 | AUC | AUC | |
| Full PHASE-T | 0.8138 | 0.8958 | 0.7413 | 0.4822 | 0.3398 | 0.8394 | 0.6369 |
| Latent-only | 0.6276 | 0.8597 | 0.5398 | 0.3271 | 0.1996 | 0.7570 | 0.6285 |
| Latent+Recon | 0.6298 | 0.8532 | 0.6405 | 0.3126 | 0.1954 | 0.7745 | 0.6156 |
| No-skip | 0.7093 | 0.7684 | 0.5904 | 0.3640 | 0.2371 | 0.7415 | 0.6354 |
| Model | Tohoku | NCNP |
|---|---|---|
| CLAP | 0.7663 | 0.5021 |
| SEEG-NET | 0.7636 | 0.5528 |
| PHASE-T, frozen | 0.7775 | 0.6292 |
| Method | Sentence onset | Speech | Volume | Pitch |
|---|---|---|---|---|
| PopT + BrainBERT | 0.90 | 0.93 | 0.87 | 0.74 |
| BaRISTA | 0.91 | 0.92 | 0.85 | 0.73 |
| MVPFormer | 0.87 | 0.90 | 0.88 | 0.83 |
| NeuroCLUS | 0.96 | 0.99 | 0.92 | 0.83 |
| PHASE-ST, frozen | 0.9176 | 0.9703 | 0.9450 | 0.8822 |
| PHASE-ST, fine-tuned | 0.9245 | 0.9786 | 0.9515 | 0.8886 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Source | Participants | Recs | Channel-hours | Sampling | Recording context | Evaluation use |
|---|---|---|---|---|---|---|
| Omni-iEEG | 151 | 399 | 5,235 | 50.67% | Clinical monitoring; official training partition | Test participants and tasks |
| Hatano | 129 | 144 | 10,315 | 43.29% | Clinical monitoring; Detroit non-REM sleep | Tohoku and NCNP cohorts |
| NYU Podcast | 8 | 8 | 632 | 2.68% | Naturalistic listening | Pretraining only |
| BrainTreebank | 10 | 17 | 5,434 | 3.36% | Naturalistic listening; BaRISTA pretraining split | 7 test recordings |
| Parameter | Setting |
|---|---|
| Waveform units and grid | Microvolts at 1 kHz; released reference convention retained |
| Pretraining crop | 60 seconds (60,000 samples) |
| High-pass filter | Second-order Butterworth, 2-Hz corner |
| Mains-notch filters | 50 and 60 Hz, each with quality factor |
| Deterministic filtering | Zero-phase product of the squared-magnitude filter responses, applied in the frequency domain after reflection padding |
| Input amplitude transform | for filtered waveforms with scale parameter , preserving absolute signal amplitudes |
| Parameter | Setting |
|---|---|
| Requested mask fraction | 55% of fine-grid tokens |
| Short mask spans | 2–4 tokens, probability 0.35 |
| Medium mask spans | 16–62 tokens, probability 0.45 |
| Long mask spans | 156–312 tokens, probability 0.20 |
| Optimizer | AdamW; , |
| Weight decay | 0.1 on matrix weights |
| Event | Window (60 s) | Anatomy | Pathol. brain region | ||||
| Model | HFO | Ictal | Sleep | Lobe-5 | Region-12 | Channel | Outcome |
| F1 | F1 | F1 | F1 | F1 | AUC | AUC | |
| LSTM+Attention | 0.7338 | – | – | – | – | – | – |
| TimesNet | 0.7652 | – | – | – | – | – | – |
| PatchTST | 0.7726 | – | – | – | – | – | – |
| eHFO | – | – | – | – | – | 0.6611 | 0.4521 |
| Event | Window (60 s) | Anatomy | Pathol. brain region | ||||
|---|---|---|---|---|---|---|---|
| Representation | HFO | Ictal | Sleep | Lobe-5 | Region-12 | Channel | Outcome |
| F1 | F1 | F1 | F1 | F1 | AUC | AUC | |
| Full PHASE-T | 0.8138 | 0.8958 | 0.7413 | 0.4822 | 0.3398 | 0.8394 | 0.6369 |
| Direct targets (no encoder) | 0.6223 ∗ | 0.8112 | 0.3089 | 0.2431 | 0.1425 | 0.7293 | 0.5962 |
| Model | Training regime | Tohoku | NCNP |
|---|---|---|---|
| TimeConv-CNN | Supervised source model | 0.7434 | 0.4716 |
| CLAP | Supervised source model | 0.7663 | 0.5021 |
| SEEG-NET | Supervised source model | 0.7636 | 0.5528 |
| PHASE-T | Frozen encoder + source probe | 0.7775 | 0.6292 |
| Model | MAYO | FNUSA |
|---|---|---|
| NeuroCLUS | 0.40 | 0.51 |
| MVPFormer | 0.36 | 0.46 |
| PopT | 0.34 | 0.43 |
| BaRISTA | 0.30 | 0.45 |
| PHASE-T | 0.7337 | 0.6656 |
| Quantity | MAYO | FNUSA |
|---|---|---|
| Training participants | 4 | 4 |
| Training clips | 21,665 | 33,805 |
| Test participants | 20 | 10 |
| Test clips | 133,517 | 159,313 |
| Binary F1 | 0.7337 | 0.6656 |
| ROC-AUC | 0.9571 | 0.8625 |
| context given to each channel | probe fitted on | pooled | mean [95%] | median |
|---|---|---|---|---|
| PHASE-T, no context | — | 0.841 | — | — |
| 1 channel | same | 0.843 | 0 | 0 |
| 8 channels, same minute | same | 0.853 | [ , ] | |
| 32 channels, same minute | same | 0.859 | [ , ] | |
| 64 channels, same minute | same | 0.862 | [ , ] | |
| same channels, another minute | another minute | 0.850 | [ , ] |