NeurDuo-EEG: A Long-Sequence EEG Foundation Model with Persistent State and Explicit Memory
Authors: Yifan Wang, Haiping Liu, Yang Cui, Wenhao Cai, Shuhang Li, Xiaoyang Huang, Xianyang Liu, Jingyu Sun, +8 more
Organizations: University of Manchester · ETH Zürich · Shanghai Jiao Tong University · Anhui University · ELLIS Institute Finland · Aalto University · Institute of Automation, Chinese Academy of Sciences · University of Chinese Academy of Sciences
Electroencephalography (EEG) is recorded continuously over hours, with relevant dynamics spanning timescales from milliseconds to hours. Most EEG foundation models nevertheless process fixed windows independently, limiting their ability to capture information encoded in long-timescale dynamics. State-space architectures enable persistent recurrent processing, but long-range information remains implicitly compressed in recurrent states. We present NeurDuo-EEG, a causal EEG foundation model with channel-resolved persistent memory. NeurDuo-EEG introduces multi-timescale memory management with learned consolidation and selective retrieval, enabling persistent modelling of continuous EEG with fixed-size state. It is pre-trained on 3,955 hours of EEG from 17 public datasets using multichannel autoregressive prediction of discrete spectral codes. Across three short-window and two long-sequence downstream tasks, NeurDuo-EEG achieves the best performance on four of five benchmarks, including all three short-window tasks and seizure detection, where AUC-PR improves from 0.285 to 0.471 over the strongest non-NeurDuo baseline. NeurDuo-EEG also remains competitive on sleep staging and supports efficient streaming inference, with nearly constant per-chunk latency as the available history grows to one hour. Notably, the Small variant achieves this with only 4.7M backbone parameters. These results demonstrate the value of persistent, multi-timescale modelling for both long-sequence and short-window EEG analysis. Our code is available at https://github.com/YifaNNW/NeurDuo-EEG.
Figures & tables
Figure 1: Overview of NeurDuo and its two-stage pre-training framework. A frozen spectral tokenizer provides discrete per-channel targets. The backbone processes coordinate-aware EEG streams through fast and slow recurrent pathways, consolidates recent states into bounded memory, and retrieves relevant past information for per-channel next-token prediction.
Model
H
df
Lf
ds
Ls
Params
Small
128
256
6
192
3
4.70M
Base
192
480
10
320
5
24.31M
Large
256
672
15
448
7
67.83M
Table 1: NeurDuo model configurations.
Methods
Model Size
FACED
KaggleERN
SEED-VIG
Bal. Acc. ↑
Cohen’s κ↑
ROC-AUC ↑
AUC-PR ↑
Pearson’s r↑
R2↑
BP-GBDT
–
0.173
0.070
0.557±0.005
0.753±0.009
0.445
0.031
FBCov-TS-Lin
–
0.185
0.084
0.613±0.023
0.778±0.023
0.344
0.110
ERP-Lin
–
0.229
0.130
0.623±0.019
0.782±0.012
0.020
−0.019
BIOT
3.2M
0.364±0.016
0.283±0.017
0.496±0.030
0.708±0.023
0.535±0.065
0.179±0.136
CBraMod
4.9M
0.537±0.008
0.476±0.010
0.571±0.015
0.755±0.009
0.529±0.038
0.196±0.051
Table 2: Short-window downstream performance. † LaBraM was pre-trained on KaggleERN; see Appendix D.1 .
Methods
Model Size
Sleep-EDF
CHB-MIT (15 ch)
CHB-MIT (2 ch)
Macro-F1 ↑
Cohen’s κ↑
AUC-PR ↑
FP ⋅ h -1 ↓
AUC-PR ↑
FP ⋅ h -1 ↓
BP-GBDT
–
0.710±0.013
0.695±0.021
0.209±0.109
3.34
0.208±0.057
3.28
FBCov-TS-Lin
–
0.674±0.018
0.666±0.019
0.285±0.136
2.14
0.224±0.124
4.12
ERP-Lin
–
0.306±0.044
0.255±0.031
0.006±0.003
15.42
0.009±0.007
12.65
BIOT †
3.2M
0.731±0.017
0.733±0.017
0.208±0.059
4.73
0.124±0.094
6.12
CBraMod
4.9M
0.759±0.009
0.749±0.014
0.163±0.085
4.48
0.277±0.150
3.40
Table 3: Long-sequence downstream performance. CHB-MIT is evaluated using either all 15 channels or only T7 and P8. † BIOT was pre-trained on SHHS and CHB-MIT; see Appendix D.1 .
Context
L
Sleep-EDF
CHB-MIT (15 ch)
CHB-MIT (2 ch)
Macro-F1 ↑
Cohen’s κ↑
AUC-PR ↑
FP ⋅ h -1 ↓
AUC-PR ↑
FP ⋅ h -1 ↓
30 s †
1
0.696±0.010
0.695±0.011
0.324±0.095
2.02
0.241±0.188
7.96
1 min
2
0.707±0.009
0.702±0.010
0.322±0.153
4.12
0.290±0.153
4.05
2 min
4
0.725±0.009
0.722±0.010
0.404±0.160
5.83
0.358±0.137
1.64
5 min
10
0.744±0.015
0.742±0.005
0.331±0.128
0.80
0.375±0.156
2.18
10 min
20
0.755±0.008
0.749±0.011
0.413±0.174
0.96
0.370±0.177
1.62
Table 4: Effect of causal context length on NeurDuo (Small). L denotes the number of consecutive 30 s units; all other evaluation settings remain fixed. †L=1 provides no cross-unit context.
Figure 2: Inference latency on CHB-MIT.
Variant
FACED
Sleep-EDF
CHB-MIT (15 ch)
Bal. Acc. ↑
κ↑
Macro-F1 ↑
κ↑
AUC-PR ↑
FP ⋅ h -1 ↓
NeurDuo (Small)
0.589±0.005
0.534±0.006
0.755±0.008
0.749±0.011
0.413±0.174
0.96
w/o per-channel tokens
0.460±0.011
0.391±0.012
0.745±0.011
0.739±0.008
0.412±0.096
1.10
w/o channel attention
0.523±0.009
0.460±0.011
0.753±0.016
0.750±0.009
0.340±0.125
4.24
w/o memory retrieval
0.506±0.012
0.442±0.013
0.744±0.012
0.736±0.010
0.389±0.112
1.08
w/o slow stream & memory
0.545±0.015
0.485±0.017
0.742±0.017
0.738±0.010
0.394±0.179
0.47
Table 5: Component ablations and temporal-order control for NeurDuo (Small). Shuffling is not applicable to FACED, whose samples are independent windows.
State
Sleep-EDF
CHB-MIT (15 ch)
Macro-F1 ↑
κ↑
AUC-PR ↑
FP ⋅ h -1 ↓
Reset
0.745±0.014
0.738±0.013
0.307±0.191
5.16
Carry
0.755±0.008
0.749±0.011
0.413±0.174
0.96
Table 6: Persistent state in NeurDuo (Small). Backbone state is reset every 30 s or carried across a 10 min sequence, keeping the context-head architecture fixed.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Domain
Datasets
#Rec.
Hours
#Ch.
#Windows
Clinical
TUSZ, TUEP, Siena, TUAR
8,007
1,811.9
17–19
395,424
Sleep (PSG)
PhysioNet-2018, ISRUC
11,339
1,882.3
6
406,517
Motor
EEGMMIDB, HGD, GAL, BCI IV-1
1,821
95.5
32–126
18,415
Attention
Cao 2019
516
81.9
30
17,652
Emotion
DEAP, DREAMER
1,694
46.2
14–32
7,367
Speech (auditory)
Broderick
380
20.6
128
4,056
Appendix
Table 7: Pre-training corpus by domain. All counts are after preprocessing. #Rec. is the number of recordings and Hours their total duration; #Ch. is the electrodes retained once channel labels are resolved to standard positions, given as a range where a domain’s corpora use different montages; #Windows is the number of 32 s pre-training windows the recordings are cut into, at 50% overlap. Table 8 gives the per-dataset breakdown.
Dataset
Domain
#Rec.
Hours
#Ch.
#Windows
PhysioNet-2018 ( Ghassemi et al., 2018 ; Goldberger et al., 2000 )
Sleep (PSG)
6,830
1,138.3
6
245,880
TUSZ ( Shah et al., 2018 )
Clinical
4,959
911.2
17–19
197,388
ISRUC ( Khalighi et al., 2016 )
Sleep (PSG)
4,509
744.0
6
160,637
TUEP ( Veloso et al., 2017 )
Clinical
2,697
631.7
17–19
138,035
Siena ( Detti et al., 2020 )
Clinical
41
169.0
19
37,963
TUAR ( Hamid et al., 2020 )
Clinical
310
100.0
19
22,038
Appendix
Table 8: Pre-training corpus, per dataset, after preprocessing. The per-dataset expansion of Table 7 , ordered by hours. #Rec. is the number of recordings and Hours their total duration; #Ch. is the electrodes retained once channel labels are resolved to standard positions; #Windows is the number of 32 s pre-training windows the recordings are cut into, at 50% overlap. Where a channel count differs from the one quoted in a release’s own publication, the table reports the signal channels the released files contain.
Hyperparameter
Stage 1 (tokenizer)
Stage 2 (backbone)
Data
Chunk
128 samples ( 0.5 s), non-overlapping
Sequence length
64 chunks ( 32 s)
64 chunks ( 32 s)
Sequence stride
64 (no overlap)
32 ( 50% overlap)
Optimisation
Optimiser
AdamW
AdamW
Peak learning rate
5×10−5
3.0/2.2/1.9×10−4 (S/B/L)
Weight decay
10−4
0.05
Appendix
Table 9: Pre-training hyperparameters. Stage 1 trains the spectral tokenizer and stage 2 the backbone, with the tokenizer frozen between them. A value spanning both columns is shared by the two stages; the Tokenizer block describes the stage-1 model itself and so is not stage-dependent. Every entry is identical across the three model sizes except the peak learning rate, given as S/B/L. “–” marks a setting that does not apply to that stage.
Hyperparameter
Value
Frontend
Chunk encoder kernels
{3,7,15} , three parallel branches
Metadata encoder width
64
Channel attention
every 2 fast layers, 4 heads
Streams
SSM state dimension
16 (fast and slow)
SSM expansion factor
2 (fast and slow)
Causal convolution kernel
4 (fast), 3 (slow)
Appendix
Table 10: Architecture constants. Identical in all three model sizes; the widths and depths that scale are given in Table 1 . One step is one 0.5 s chunk, so the durations in parentheses follow from the step counts beside them.
Figure 3: Pre-training curves for the three sizes of Table 1 , on the held-out recordings. Small and Base reach their minima as the schedule ends, at 6.0503 and 6.0272 nats; Large reaches its minimum of 6.0491 after 51,000 updates and then rises to 6.0715 .
Field
Used
Reason
Electrode coordinates
✓
Physical position; the only cue distinguishing channels.
Reference scheme
✓
Eight values across the pre-training corpus, reflecting how each recording was referenced.
Channel type
—
Constant: all 438,323 channels share one value.
Sampling rate
—
Constant: every recording is resampled to 256 Hz.
Device identifier
—
Near-constant: 22,740 of 22,780 recordings share one value.
Corpus identifier
—
Informative but unavailable at transfer time.
Appendix
Table 11: Metadata fields available in the pipeline and their disposition.
Absolute CE ( ↓ )
Improvement over identity ( ↑ ) for K=
Predict ahead
marginal
identity
4 s
8 s
16 s
32 s
64 s
128 s
Best
1 s
3.4435
3.2911
0.1048
0.1090
0.1097
0.1070
0.1038
0.0982
16 s
10 s
3.4418
3.2918
0.0892
0.0952
0.0988
0.1001
0.0984
0.0948
32 s
30 s
3.4433
3.2922
0.0761
0.0835
0.0891
0.0915
0.0920
0.0895
64 s
60 s
3.4398
3.2913
0.0695
0.0756
0.0817
0.0854
0.0865
0.0854
64 s
Appendix
Table 12: Exploratory history-probe results conditioned on recording-level statistics. Rows are prediction horizons and columns the length of history supplied. Marginal is a predictor given no input and Identity one given only a leave-one-out estimate of the recording’s own code distribution; the first two columns are their absolute held-out cross-entropies in nats, and the remaining six are CE(\textscIdentity)−CE(\textscIdentity+K) , the reduction in cross-entropy from adding K seconds of history to the same recording-level estimate. Best denotes the history length K with the largest improvement, also marked in bold . Measured on 19,200 windows drawn from 4,600 pre-training recordings, of which 17,026 windows support the leave-one-out estimate, split by recording.
Dataset
Task
Unit
#Ch.
Rate
L
K
Selection
Split
FACED ( Chen et al., 2023 )
Emotion, 9 -class
10 s
30
250
1
19,200
κ
subject, fixed
KaggleERN ( Mattout et al., 2014 ; Margaux et al., 2012 )
ERN, binary
2 s
19
200
1
3,840
ROC-AUC
subj., 4 folds †
SEED-VIG ( Zheng and Lu, 2017 )
Vigilance, regression
8 s
17
200
1
19,200
R2
session, fixed
Sleep-EDF ( Kemp et al., 2000 ; Goldberger et al., 2000 )
Sleep stage, 5 -class
30 s
2
100
20
12,000
κ
subject, 5 -fold
CHB-MIT ( Goldberger et al., 2000 ; Shoeb, 2009 )
Seizure, binary
30 s
15
256
20
90,000
AUC-PR
case, 5 -fold ‡
Appendix
Table 13: The five downstream tasks. The three above the rule are window tasks and the two below are sequence tasks. Unit is the interval one label is defined on, #Ch. the electrodes used and Rate their sampling rate in Hz; L is the number of consecutive units per sample, so L=1 is a single labelled segment; K is the readout budget, fixed per dataset and identical across models (Appendix D.2.1 ); Selection is the validation criterion a checkpoint is chosen on, and Split how the data is partitioned. † Not a cross-validation (Appendix C.2 ). ‡ Folds are over the 24 case folders, which come from 23 subjects; see Appendix C.5
Baseline
Architecture / feature family
#Params
Pre-training
BIOT
Linear-attention transformer over per-channel STFT patches
3.2 M
MGH, SHHS, TUAB, TUEV, CHB-MIT, IIIC
CBraMod
Criss-cross transformer over the channel × time grid
4.9 M
TUEG, >9,000 h
LaBraM-Base
Vector-quantised tokenizer with masked-token prediction
5.8 M
16 corpora, >2,500 h
EEGPT-Large
ViT with dual mask-reconstruction and alignment
25.3 M
PhysioNet-MI, TSUBenchmark, M3CV, SEED
ST-EEGFormer-S
ViT with masked autoencoding on raw EEG
25.4 M
>8 M segments; MI, P300, SSVEP
REVE-Base
ViT with 4D Fourier encoding of electrode geometry
69.2 M
92 datasets, >60,000 h, ∼25,000 subjects
Appendix
Table 14: The nine baselines. The six above the rule are pre-trained foundation models and the three below are decoders that were never pre-trained. #Params is the pre-trained backbone as loaded, excluding pre-training-only components such as decoders and projection heads; a classical pipeline has no separable backbone and readout, so the column is left blank. Pre-training reproduces each release’s own account of its corpus, which is why the entries are not stated on a common footing.
Model
FACED
KaggleERN
SEED-VIG
Sleep-EDF
CHB-MIT
BIOT
∘
∙
CBraMod
∘
LaBraM-Base
∘
∙
∘
EEGPT-Large
∘
ST-EEGFormer-S
REVE-Base
∘
∘
Appendix
Table 15: Pre-training overlap by model and task. Every model’s pre-training corpus, ours included, was audited against every downstream dataset. ∙ marks corpus overlap, the downstream corpus itself appearing in that model’s pre-training set; it occurs twice and both cases are discussed below. ∘ marks domain overlap, a different corpus of the same recording paradigm and population, which is the ordinary condition of pre-training at scale rather than leakage. Blank means neither. The three classical pipelines of Table 14 have no row here, never having been pre-trained.
Dataset
K
Determined by
Head size
Spread
FACED
19,200
1,920 dim/s at 10 s
4.92 – 4.98 M
1.012×
KaggleERN
3,840
1,920 dim/s at 2 s
1.00 – 1.05 M
1.045×
SEED-VIG
19,200
carried over from FACED
4.93 – 5.02 M
1.020×
Sleep-EDF
12,000
no-clamp bound minmNmDm
4.75 – 4.94 M
1.041×
CHB-MIT (15 Ch)
90,000
no-clamp bound minmNmDm
24.81 – 24.95 M
1.006×
CHB-MIT (2 Ch)
12,000
no-clamp bound minmNmDm
4.75 – 4.94 M
1.041×
Appendix
Table 16: Readout budget per dataset and the head sizes it produces. K=Ndtok is the product of token count and per-token width, held fixed per dataset so that head size does not track whichever tokenizer emits the most tokens. Determined by names the rule that sets it: either a fixed readout capacity per second of EEG, or the no-clamp bound minmNmDm where holding that capacity would push some model’s dtok above its width Dm . Head size is the range over the nine models that receive a shared head and Spread their ratio; the Sleep-EDF and CHB-MIT figures include the context head, which is identical across models and therefore compresses the ratio.
Activation
d/dx at x=−8
Seeds escaping
Balanced accuracy
GELU
0 (and <0 for x<−0.75 )
0/5
0.111±0.000
ReLU
0
1/5
0.165±0.108
ELU
3.4×10−4
3/5
0.219±0.124
LeakyReLU (0.01)
1.0×10−2
5/5
0.354±0.013
Appendix
Table 17: Escape from the class-prior solution on FACED with the BIOT backbone. Five seeds per activation, warm-up disabled, everything else held fixed. A run that fails to escape drives all 256 hidden units of the shared head negative and emits the training class prior thereafter, at which point balanced accuracy is exactly 1/9 , the chance level for nine classes. Rows are ordered by how much gradient the activation leaves on the negative side, evaluated at x=−8 , a value typical of the collapsed units. The recipe of the main tables adds three warm-up epochs, which is why the LeakyReLU row reads 0.354 here against the 0.364 of Table 2 .
Figure 4: Effect of pre-training. Downstream performance of NeurDuo (Small) with and without pre-training. Arrows indicate whether higher or lower values are better.
Model
W
N1
N2
N3
REM
BIOT †
0.924
0.463
0.828
0.676
0.764
CBraMod
0.914
0.458
0.845
0.805
0.775
LaBraM
0.928
0.478
0.851
0.799
0.773
EEGPT
0.922
0.455
0.829
0.766
0.725
ST-EEGFormer-S
0.923
0.413
0.826
0.672
0.740
REVE-Base
0.930
0.478
0.835
0.680
0.784
Appendix
Table 18: Per-class F1 on Sleep-EDF , five-fold means. Classes are the five AASM stages, W for wake and N1 to N3 for non-REM depth. Bold marks the best value in a column among the seven models above the rule, each in the configuration Table 3 reports. Below the rule are that same model with its state reset at every unit boundary and the difference between the two, neither of which is part of that comparison. † Domain overlap on Sleep-EDF, see Appendix D.1 .
Metric
Fold 1
Fold 2
Fold 3
Fold 4
Fold 5
Mean
t
Positive
Macro-F1
+0.0277
+0.0159
+0.0039
+0.0002
+0.0037
+0.0103
+2.02
5/5
Accuracy
+0.0089
+0.0146
+0.0044
+0.0083
+0.0043
+0.0081
+4.27
5/5
Cohen’s κ
+0.0141
+0.0189
+0.0051
+0.0107
+0.0040
+0.0106
+3.80
5/5
Appendix
Table 19: Carried minus reset on Sleep-EDF, fold by fold. Same folds, same seed, same readout and context head; only the backbone state differs. t is the paired statistic over the five differences, for which the two-sided 5% critical value at four degrees of freedom is 2.776 .
Figure 5: The two streams run at different rates. (a) Nominal time constant τ=1/(Δ∣A∣) , with Δ taken from the trained bias alone, of every state dimension of the pre-trained Small model, converted to seconds at each stream’s own step; dashed lines are medians. (b) Relative state change per second on one ten-minute CHB-MIT test sequence, averaged over the 15 channels; the shaded band marks its three seizure units.
Figure 6: Stateful streaming inference. Each 0.5 s chunk advances the fast stream by one step; every fourth chunk writes a memory token that advances the slow stream and enters the 16 s bank the fast state queries. Nothing is reset within a recording.
Model
Sleep-EDF (2 ch)
CHB-MIT (15 ch)
30 s
300 s
600 s
1,800 s
3,600 s
GB
30 s
300 s
600 s
1,800 s
3,600 s
GB
EEGPT-Large
4.0
27.3
49.8
137.6
274.0
1.22
8.5
66.0
128.2
382.9
765.5
3.46
ST-EEGFormer-S
3.9
40.5
136.6
1,221
5,119
0.97
28.9
1,934
8,061
74,302
– ∗
5.64
BIOT
6.0
23.3
40.5
106.9
209.2
0.89
6.0
23.3
42.7
107.0
209.5
0.98
LaBraM
9.9
98.0
196.4
587.2
1,174
0.05
10.0
98.1
196.2
585.9
1,177
0.20
CBraMod
5.8
5.9
6.1
15.1
38.1
0.16
5.8
9.2
19.0
93.6
281.4
0.99
Appendix
Table 20: Cost of conditioning on longer history. Latency (ms) to encode one window, with peak memory (GB) measured at 1,800 s. NeurDuo latency is measured per 0.5 s chunk after streaming the stated history.
Component
Shape
Values
KB (fp32)
Fast stream
6×(512×16SSM+512×4conv)
61,440
240.0
Slow stream
3×(384×16SSM+384×3conv)
21,888
85.5
Memory bank
8×192
1,536
6.0
Writer queue
16×256
4,096
16.0
Last output
256+192
448
1.8
Total per channel
89,408
349.2
Appendix
Table 21: The persistent state , per channel, for NeurDuo (Small). Sizes follow from the architecture constants of Table 10 and are the same at every point in a recording. The writer queue is a ring buffer of 16 slots, of which each write pools the most recent w=8 ; an implementation that kept only those eight would save 8 KB per channel.
Sleep-EDF montage, two channels
CHB-MIT montage, fifteen channels
History
ms
× real time
GB peak
MB state
ms
× real time
GB peak
MB state
30 s
4.89
102×
0.047
0.68
4.90
102×
0.061
5.12
1 min
4.89
102×
0.047
0.68
4.89
102×
0.061
5.12
5 min
4.90
102×
0.047
0.68
4.87
103×
0.061
5.12
10 min
4.84
103×
0.047
0.68
4.85
103×
0.061
5.12
30 min
4.85
103×
0.047
0.68
4.87
103×
0.061
5.12
Appendix
Table 22: Streaming a recording from thirty seconds to an hour , NeurDuo (Small). Each row advances the state to the stated amount of history and then measures the next chunks, averaged over a full write cycle. Nothing in the table depends on the history, which is the property § 3.2.2 claims and this table exists to check.
Electroencephalography (EEG) is a critical, non-invasive method to monitor electrical brain activity. EEGs can span anywhere from a couple seconds to multiple hours, posing a major hurdle for existing deep learning methods due to two major factors: (1) existing EEG models are predominantly built upon the attention mechanism, incurring quadratic scaling as the sequence length increases, and (2) raw EEG signals must be processed in a sliding-window fashion due to fixed-length input requirements, preventing global understanding of the entire signal. To this extent, we propose CaMBRAIN - the first Causal, Mamba-based state space model (SSM) capable of real-time inference of EEG signals, arguing that bidirectional approaches are needlessly expensive given the causal, unidirectional nature of EEG. However, training such a model is non-trivial, as crucial EEG events can be extremely brief - within fractions of a second - yet separated by long intervals spanning minutes. Current EEG methods use self-supervised objectives that optimize for signal reconstruction, but these are not well suited for streaming SSMs; they fail to explicitly train the hidden state to retain the salient long-range context needed for streaming inference. We therefore introduce a multi-stage self-supervised training pipeline specifically tailored to encourage long-range memory retention and strong performance on EEG signals, while preserving the linear-time complexity of state space models. CaMBRAIN achieves state-of-the-art (SOTA) results across 3 different EEG datasets with >10x higher throughput than existing models, enabling the first model capable of long-range, continuous inference of variable-length EEG signals.
Abhilash Durgam, Nyle Siddiqui, Jeffrey A. Chan-Santiago +3
CRCV, University of Central Florida · Department of MAE, University of Central Florida · Department of Neurology, Loma Linda University
Electroencephalography (EEG) foundation models aim to learn reusable representations from large-scale unlabeled recordings. A common pretraining strategy is masked waveform reconstruction, but applying supervision directly to noisy EEG may encourage models to recover predictable background activity, acquisition effects, and artifacts rather than neural structure that transfers across tasks. This raises a central question: what should an EEG foundation model predict to learn transferable representations? We introduce EEG-JEPA a structured latent-prediction framework for EEG foundation modeling. Rather than reconstructing masked voltage samples, a masked context encoder and predictor infer contextual latent states produced by an exponential-moving-average target encoder that observes the complete input. EEG-JEPA organizes target design along three complementary dimensions: target content specifies what representation is predicted, target support specifies where prediction occurs over structured electrode--time regions through Neurotopology-Aware Multi-scale Electrode-Temporal Masking (N-MET), and target depth specifies at which encoder layers supervision is applied. Together, these designs shift EEG pretraining from recovering missing measurements to inferring latent states from structured electrode--time context. We evaluate EEG-JEPA through controlled objective comparisons, frozen multitask transfer, and full fine-tuning. Under the same backbone, pretraining corpus, and training duration, EEG-JEPA improves the 14-task frozen macro balanced accuracy from 40.49% to 50.42% over CBraMod-style masked waveform reconstruction. Multi-source continuation further raises this result to 52.94%, the highest average among the EEG foundation models evaluated on EEG-FM-Bench. Under protocol-matched full fine-tuning, EEG-JEPA also improves the nine-task average balanced accuracy from 68.98% to 70.65%.
Jinhao Li, Zhiyuan Ma, Xueqiao Han +8
Tsinghua Laboratory of Brain and Intelligence, Tsinghua University · School of Basic Medical Sciences, Tsinghua Medicine, Tsinghua University · School of Biomedical Engineering, Tsinghua Medicine, Tsinghua University +2
Foundation models offer a promising paradigm for Electroencephalography (EEG) analysis, leveraging generalizable representations from vast unlabeled datasets. Yet, Transformer-based architectures face a critical bottleneck: global attention mechanisms couple the attention memory state to the signal duration, causing memory overflow during continuous monitoring. To address this, we introduce S-CEReBrO (Streaming CEReBrO), an evolution of the CEReBrO architecture designed for continuous monitoring. Our novel Windowed Alternating Attention mechanism factorizes attention computation into fixed-size spatiotemporal windows, so that under streaming, only the active window remains resident and the attention state is bounded independently of signal duration. Empirical scaling analysis shows that windowed alternating attention can process signals 100X longer than full self-attention and 3X longer than low-rank linear attention. Compared to low-rank linear attention on long contexts, windowed alternating attention requires 55% of the memory while increasing inference throughput by 2.1X. Pre-trained on >25,000 hours of recordings from >12,000 subjects, S-CEReBrO achieves state-of-the-art performance on 7 of 11 downstream tasks, with up to 60% fewer parameters. This work represents a significant step toward the realization of efficient, generalizable, and continuous EEG monitoring. An accompanying code repository is available. An accompanying code repository is available.
Glenn Anta Bucagu, Thorir Mar Ingolfsson, Yawei Li +1
ETH Zurich, Zurich, Switzerland · Nanyang Technological University, Singapore · University of Bologna, Bologna, Italy