EEGDM: Label-Efficient EEG Representation Learning with Generative Diffusion Model
Authors: Jia Hong Puah, Sim Kuan Goh, Ziwei Zhang, Zixuan Ye, Chow Khuen Chan, Kheng Seang Lim, Si Lei Fong, Kok Sin Woon, +1 more
Organizations: School of Artificial Intelligence and Robotics, Xiamen University Malaysia · School of Energy and Chemical Engineering, Xiamen University Malaysia · Department of Biomedical Engineering, Universiti Malaya · Department of Medicine, Universiti Malaya · Thrust of Carbon Neutrality and Climate Change, Hong Kong University of Science and Technology (Guangzhou) · College of Computing & Data Science, Nanyang Technological University
Electroencephalography (EEG) is a critical tool for monitoring brain activity and diagnosing neurological disorders such as epilepsy. However, learning meaningful representations from raw EEG signals remains challenging due to limited annotations, substantial inter-subject variability, and complex temporal dynamics. Recent EEG foundation models (FMs) have demonstrated promising performance through transformer-based architectures and large-scale self-supervised pretraining, yet they often incur substantial computational costs and exhibit diminishing returns with increasing model and dataset scale, limiting their practicality in clinical settings. To address these challenges, we propose EEGDM, a diffusion-based EEG representation learning framework. Specifically, EEGDM introduces a Structured State-Space Model for Diffusion Pretraining (SSMDP) that effectively captures long-range temporal dependencies through generative diffusion training on unlabeled EEG data. The representations are subsequently leveraged for downstream tasks via our Latent Fusion Module (LFM), which integrates multi-layer latent features of SSMDP. We evaluate EEGDM on three EEG benchmarks spanning EEG event classification (TUEV), seizure detection (CHB-MIT), and seizure classification (IIIC). Compared with existing state-of-the-art methods, including EEG FMs, EEGDM achieves competitive performance across diverse tasks and datasets, exceeding existing methods in most cases while requiring substantially fewer samples for both pretraining and downstream adaptation. These results demonstrate that diffusion-based learning can unlock more discriminative EEG representations, with direct implications for epilepsy diagnosis and management. Our source code and pretrained checkpoints are publicly available at: https://github.com/jhpuah/EEGDM.
Figures & tables
Fig. 1: EEGDM builds on a structured state-space model (SSMDP) learned via generative diffusion-based pretraining and finetunes a Latent Fusion Module (LFM) for downstream tasks. It generalizes across in-domain EEG event classification as well as cross-task and cross-domain settings, including seizure detection and seizure classification under diverse EEG configurations. Experimental results demonstrate that EEGDM is highly competitive with recent EEG foundation models, despite being trained on substantially less data for representation learning.
Fig. 2: Illustration of the proposed SSMDP for representation learning. The model learned to denoise and reverse a noised EEG signal by predicting the diffusion velocity, conditioned on diffusion step embedding and channel embedding. SSMDP built upon DiffWave and S4D architectures, and trained using DDPM framework. Subsequently, the latent activities from the gate or filter channels were then leveraged for downstream classification.
Fig. 3: From the pretrained SSMDP, we extract the latent activities of each EEG channel as a four-dimensional tensor: (# channels, # layers, # time points, # gate channels per layer). The latent activities are subsequently pooled along the temporal dimension to form compact latent tokens, reducing the representation to (# channels, # layers, # pools, # gate channels per layer). Different colors denote different EEG channels.
Fig. 4: Architecture of the proposed LFM. LFM comprises two main modules: LFM-CL and LFM-T. Left: The LFM-CL processes tokens from one pool at a time, with each block handling tokens from a single layer. Trainable fusion tokens are passed sequentially between blocks, integrating nonisomorphic representations from different layers via cross-attention. Middle: To capture time-invariant representations, the same fusion model is applied to every pool. Right: The LFM-T processes the fused representations from all pools. Subsequently, global average pooling aggregates the learned representations, followed by a linear classification head that produces the final predictions.
Model
Size
Kappa
BAcc
WF1
Training from Scratch (non-FM)
SPaRCNet [ 29 ]
0.79M
42.33 ± 1.81
41.61 ± 2.62
70.24 ± 1.04
ContraWR [ 64 ]
1.6M
39.12 ± 2.37
43.84 ± 3.49
68.93 ± 1.36
FFCL [ 65 ]
2.4M
37.32 ± 1.88
39.79 ± 1.04
67.83 ± 1.20
CNN-Trans [ 66 ]
3.2M
38.15 ± 1.34
40.87 ± 1.61
68.54 ± 2.93
ST-Trans [ 67 ]
3.5M
37.65 ± 3.06
39.84 ± 2.28
68.23 ± 1.90
TABLE I: Comparative analysis (%) on Multi-event TUEV with FMs and non-FMs.
Model Setting
Kappa
BAcc
WF1
EEGDM
74.23 ± 1.36
75.57 ± 1.47
86.88 ± 0.66
- μ -law Signal Embedding
72.50 ± 3.02
69.12 ± 1.81
85.79 ± 1.41
- Cosine Noise Schedule
70.97 ± 1.04
71.19 ± 0.92
85.07 ± 0.49
- Noiseless Forward Process
71.39 ± 1.46
56.19 ± 3.34
84.87 ± 0.54
- Bidirectional SSM
65.92 ± 0.95
65.57 ± 2.01
82.31 ± 1.07
TABLE II: Ablations of μ -law, noise scheduler, noiseless forward process, and bidirectional SSM (%).
Step
Kappa
BAcc
WF1
No
74.23 ± 1.36
75.57 ± 1.47
86.88 ± 0.66
1
74.46 ± 1.39
74.87 ± 1.83
86.94 ± 0.68
2
73.73 ± 1.16
74.83 ± 1.55
86.66 ± 0.55
3
69.05 ± 1.63
69.89 ± 1.49
84.34 ± 0.81
TABLE III: Impacts of different diffusion time step (%).
Pooling
Latent
Kappa
BAcc
WF1
Std
Gate
74.23 ± 1.36
75.57 ± 1.47
86.88 ± 0.66
Std
Filter
73.46 ± 2.33
68.74 ± 0.63
86.49 ± 1.18
Average
Gate
71.66 ± 1.13
57.71 ± 1.89
85.24 ± 0.57
Average
Filter
69.83 ± 1.92
68.33 ± 3.17
84.46 ± 1.05
TABLE IV: Impacts of pooling strategies and latent activities (%).
Layers
Size
Kappa
BAcc
WF1
All
6.9M
74.23 ± 1.36
75.57 ± 1.47
86.88 ± 0.66
First half
7.0M
62.97 ± 2.15
64.57 ± 1.54
81.00 ± 1.15
Second half
7.0M
71.60 ± 1.14
68.40 ± 1.92
85.60 ± 0.52
First quarter
6.9M
59.63 ± 2.49
58.75 ± 3.13
78.88 ± 1.43
Second quarter
6.9M
68.40 ± 2.66
69.47 ± 1.45
83.80 ± 1.35
Third quarter
6.9M
68.32 ± 1.22
64.39 ± 1.37
83.96 ± 0.63
TABLE V: Impacts of using latent activities from different layers (%).
Fusion
Size
Kappa
BAcc
WF1
Base
6.9M
74.23 ± 1.36
75.57 ± 1.47
86.88 ± 0.66
No
7.0M
61.60 ± 3.69
51.13 ± 7.15
79.93 ± 2.26
Mean
7.0M
65.05 ± 1.41
59.40 ± 8.67
82.08 ± 0.83
TABLE VI: Impacts of different latent fusion strategies (%).
Fig. 5: Visualization of SSMDP-generated EEG traces.
Model
Size
BAcc
AUC-PR
AUROC
Training from Scratch on CHB-MIT (non-FM)
SPaRCNet [ 29 ]
0.79M
58.76 ± 1.91
12.47 ± 1.19
81.43 ± 1.48
ContraWR [ 64 ]
1.6M
63.44 ± 0.02
22.64 ± 1.74
80.97 ± 1.14
FFCL [ 65 ]
2.4M
62.62 ± 1.04
20.49 ± 3.46
82.71 ± 0.51
CNN-Trans [ 66 ]
3.2M
63.89 ± 0.67
24.79 ± 2.27
86.62 ± 0.82
ST-Trans [ 67 ]
3.5M
59.15 ± 1.95
14.22 ± 0.94
82.37 ± 4.91
TABLE VII: Generalization (%) of SSMDP pretrained on TUEV (with interictal but no labeled ictal seizures) to Cross Task & Domain Downstream CHB-MIT (with pediatric seizures)
Train
Test
Kappa
BAcc
WF1
200 Hz
200 Hz
74.23 ± 1.36
75.57 ± 1.47
86.88 ± 0.66
200 Hz
190 Hz
70.26 ± 2.00
72.78 ± 1.38
85.17 ± 1.02
200 Hz
180 Hz
66.54 ± 1.60
68.36 ± 2.07
83.38 ± 0.81
TABLE VIII: Generalization (%) EEGDM to different sampling rates (TUEV).
Pretraining
Kappa
BAcc
WF1
Diffusion Pretraining
74.23 ± 1.36
75.57 ± 1.47
86.88 ± 0.66
Masked Reconstruction
59.69 ± 2.04
54.98 ± 2.54
78.81 ± 1.24
TABLE IX: Ablation of Diffusion Pretraining on TUEV (%).
Model
BAcc
AUC-PR
AUROC
Finetuned with 100% Training Data
BIOT [ 1 ]
70.68 ± 4.57
32.77 ± 4.60
87.61 ± 2.84
LaBraM-Base [ 2 ]
70.75 ± 3.58
32.87 ± 4.02
86.79 ± 1.99
CBraMod [ 9 ]
73.98 ± 2.84
36.89 ± 3.82
88.92 ± 1.54
Finetuned on smaller subsets of Training Data
EEGDM (Ours)
85.82 ± 1.04
48.45 ± 0.48
89.26 ± 0.86
TABLE X: Generalization of EEGDM to CHB-MIT dataset and its label efficiency (%).
Model
Kappa
BAcc
WF1
5 Seeds (Seed 0 ∼ 4)
BIOT [ 1 ]
36.29 ± 1.75
49.06 ± 1.44
48.18 ± 1.49
CBraMod [ 9 ]
40.77 ± 0.25
49.33 ± 0.29
51.02 ± 0.30
CodeBrain [ 69 ]
28.25 ± 0.70
38.74 ± 1.35
40.28 ± 0.64
EEGMamba [ 16 ]
44.03 ± 0.14
52.84 ± 0.23
54.64 ± 0.13
Gram [ 10 ]
50.48 ± 2.21
59.37 ± 2.16
60.20 ± 1.74
TABLE XI: Comparative results (%) on the IIIC seizure classification dataset. Statistical tests are conducted between the top two methods. ∗/∗∗ and †/†† denote statistical significance at p<0.01 and p<0.001 for the Wilcoxon signed-rank test and paired t -test, respectively.
Component
Value
Kappa
BAcc
WF1
8
71.40 ± 1.16
74.23 ± 1.80
85.60 ± 0.63
# of fusion tokens
16*
74.23 ± 1.36
75.57 ± 1.47
86.88 ± 0.66
24
71.52 ± 1.65
73.03 ± 2.55
85.55 ± 0.85
4
70.74 ± 1.68
73.05 ± 1.93
85.04 ± 0.91
# of encoder blocks
8*
74.23 ± 1.36
75.57 ± 1.47
86.88 ± 0.66
12
72.71 ± 1.52
74.40 ± 1.17
86.20 ± 0.75
TABLE XII: Hyperparameter sensitivity analysis of the LFM (%). * indicate the default value.
Department of Electrical and Computer Engineering, Lebanese American University, Byblos, Lebanon · Institute of Applied Artificial Intelligence, T´ELUQ University, Montreal, Canada