NeurDuo-EEG: A Long-Sequence EEG Foundation Model with Persistent State and Explicit Memory
Organizations: University of Manchester · ETH Zürich · Shanghai Jiao Tong University · Anhui University · ELLIS Institute Finland · Aalto University · Institute of Automation, Chinese Academy of Sciences · University of Chinese Academy of Sciences
Abstract
Electroencephalography (EEG) is recorded continuously over hours, with relevant dynamics spanning timescales from milliseconds to hours. Most EEG foundation models nevertheless process fixed windows independently, limiting their ability to capture information encoded in long-timescale dynamics. State-space architectures enable persistent recurrent processing, but long-range information remains implicitly compressed in recurrent states. We present NeurDuo-EEG, a causal EEG foundation model with channel-resolved persistent memory. NeurDuo-EEG introduces multi-timescale memory management with learned consolidation and selective retrieval, enabling persistent modelling of continuous EEG with fixed-size state. It is pre-trained on 3,955 hours of EEG from 17 public datasets using multichannel autoregressive prediction of discrete spectral codes. Across three short-window and two long-sequence downstream tasks, NeurDuo-EEG achieves the best performance on four of five benchmarks, including all three short-window tasks and seizure detection, where AUC-PR improves from to over the strongest non-NeurDuo baseline. NeurDuo-EEG also remains competitive on sleep staging and supports efficient streaming inference, with nearly constant per-chunk latency as the available history grows to one hour. Notably, the Small variant achieves this with only 4.7M backbone parameters. These results demonstrate the value of persistent, multi-timescale modelling for both long-sequence and short-window EEG analysis. Our code is available at https://github.com/YifaNNW/NeurDuo-EEG.
Figures & tables
| Model | Params | |||||
| Small | 128 | 256 | 6 | 192 | 3 | 4.70M |
| Base | 192 | 480 | 10 | 320 | 5 | 24.31M |
| Large | 256 | 672 | 15 | 448 | 7 | 67.83M |
| Methods | Model Size | FACED | KaggleERN | SEED-VIG | |||
| Bal. Acc. | Cohen’s | ROC-AUC | AUC-PR | Pearson’s | |||
| BP-GBDT | – | 0.173 | 0.070 | 0.445 | 0.031 | ||
| FBCov-TS-Lin | – | 0.185 | 0.084 | 0.344 | 0.110 | ||
| ERP-Lin | – | 0.229 | 0.130 | 0.020 | |||
| BIOT | 3.2M | ||||||
| CBraMod | 4.9M | ||||||
| Methods | Model Size | Sleep-EDF | CHB-MIT (15 ch) | CHB-MIT (2 ch) | |||
| Macro-F1 | Cohen’s | AUC-PR | FP h -1 | AUC-PR | FP h -1 | ||
| BP-GBDT | – | ||||||
| FBCov-TS-Lin | – | ||||||
| ERP-Lin | – | ||||||
| BIOT † | 3.2M | ||||||
| CBraMod | 4.9M | ||||||
| Context | Sleep-EDF | CHB-MIT (15 ch) | CHB-MIT (2 ch) | ||||
| Macro-F1 | Cohen’s | AUC-PR | FP h -1 | AUC-PR | FP h -1 | ||
| s † | 1 | ||||||
| min | 2 | ||||||
| min | 4 | ||||||
| min | 10 | ||||||
| min | 20 | ||||||
| Variant | FACED | Sleep-EDF | CHB-MIT (15 ch) | |||
| Bal. Acc. | Macro-F1 | AUC-PR | FP h -1 | |||
| NeurDuo (Small) | ||||||
| w/o per-channel tokens | ||||||
| w/o channel attention | ||||||
| w/o memory retrieval | ||||||
| w/o slow stream & memory | ||||||
| State | Sleep-EDF | CHB-MIT (15 ch) | ||
| Macro-F1 | AUC-PR | FP h -1 | ||
| Reset | ||||
| Carry | ||||
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Domain | Datasets | #Rec. | Hours | #Ch. | #Windows |
| Clinical | TUSZ, TUEP, Siena, TUAR | 8,007 | 1,811.9 | 17–19 | 395,424 |
| Sleep (PSG) | PhysioNet-2018, ISRUC | 11,339 | 1,882.3 | 6 | 406,517 |
| Motor | EEGMMIDB, HGD, GAL, BCI IV-1 | 1,821 | 95.5 | 32–126 | 18,415 |
| Attention | Cao 2019 | 516 | 81.9 | 30 | 17,652 |
| Emotion | DEAP, DREAMER | 1,694 | 46.2 | 14–32 | 7,367 |
| Speech (auditory) | Broderick | 380 | 20.6 | 128 | 4,056 |
| Dataset | Domain | #Rec. | Hours | #Ch. | #Windows |
| PhysioNet-2018 ( Ghassemi et al., 2018 ; Goldberger et al., 2000 ) | Sleep (PSG) | 6,830 | 1,138.3 | 6 | 245,880 |
| TUSZ ( Shah et al., 2018 ) | Clinical | 4,959 | 911.2 | 17–19 | 197,388 |
| ISRUC ( Khalighi et al., 2016 ) | Sleep (PSG) | 4,509 | 744.0 | 6 | 160,637 |
| TUEP ( Veloso et al., 2017 ) | Clinical | 2,697 | 631.7 | 17–19 | 138,035 |
| Siena ( Detti et al., 2020 ) | Clinical | 41 | 169.0 | 19 | 37,963 |
| TUAR ( Hamid et al., 2020 ) | Clinical | 310 | 100.0 | 19 | 22,038 |
| Hyperparameter | Stage 1 (tokenizer) | Stage 2 (backbone) | |
| Data | Chunk | samples ( s), non-overlapping | |
| Sequence length | chunks ( s) | chunks ( s) | |
| Sequence stride | (no overlap) | ( overlap) | |
| Optimisation | Optimiser | AdamW | AdamW |
| Peak learning rate | (S/B/L) | ||
| Weight decay | |||
| Hyperparameter | Value | |
| Frontend | Chunk encoder kernels | , three parallel branches |
| Metadata encoder width | ||
| Channel attention | every fast layers, heads | |
| Streams | SSM state dimension | (fast and slow) |
| SSM expansion factor | (fast and slow) | |
| Causal convolution kernel | (fast), (slow) |
| Field | Used | Reason |
| Electrode coordinates | ✓ | Physical position; the only cue distinguishing channels. |
| Reference scheme | ✓ | Eight values across the pre-training corpus, reflecting how each recording was referenced. |
| Channel type | — | Constant: all channels share one value. |
| Sampling rate | — | Constant: every recording is resampled to Hz. |
| Device identifier | — | Near-constant: of recordings share one value. |
| Corpus identifier | — | Informative but unavailable at transfer time. |
| Absolute CE ( ) | Improvement over identity ( ) for | ||||||||
| Predict ahead | marginal | identity | s | s | s | s | s | s | Best |
| s | 3.4435 | 3.2911 | 0.1048 | 0.1090 | 0.1097 | 0.1070 | 0.1038 | 0.0982 | s |
| s | 3.4418 | 3.2918 | 0.0892 | 0.0952 | 0.0988 | 0.1001 | 0.0984 | 0.0948 | s |
| s | 3.4433 | 3.2922 | 0.0761 | 0.0835 | 0.0891 | 0.0915 | 0.0920 | 0.0895 | s |
| s | 3.4398 | 3.2913 | 0.0695 | 0.0756 | 0.0817 | 0.0854 | 0.0865 | 0.0854 | s |
| Dataset | Task | Unit | #Ch. | Rate | Selection | Split | ||
| FACED ( Chen et al., 2023 ) | Emotion, -class | s | subject, fixed | |||||
| KaggleERN ( Mattout et al., 2014 ; Margaux et al., 2012 ) | ERN, binary | s | ROC-AUC | subj., folds † | ||||
| SEED-VIG ( Zheng and Lu, 2017 ) | Vigilance, regression | s | session, fixed | |||||
| Sleep-EDF ( Kemp et al., 2000 ; Goldberger et al., 2000 ) | Sleep stage, -class | s | subject, -fold | |||||
| CHB-MIT ( Goldberger et al., 2000 ; Shoeb, 2009 ) | Seizure, binary | s | AUC-PR | case, -fold ‡ |
| Baseline | Architecture / feature family | #Params | Pre-training |
| BIOT | Linear-attention transformer over per-channel STFT patches | M | MGH, SHHS, TUAB, TUEV, CHB-MIT, IIIC |
| CBraMod | Criss-cross transformer over the channel time grid | M | TUEG, h |
| LaBraM-Base | Vector-quantised tokenizer with masked-token prediction | M | corpora, h |
| EEGPT-Large | ViT with dual mask-reconstruction and alignment | M | PhysioNet-MI, TSUBenchmark, M3CV, SEED |
| ST-EEGFormer-S | ViT with masked autoencoding on raw EEG | M | M segments; MI, P300, SSVEP |
| REVE-Base | ViT with 4D Fourier encoding of electrode geometry | M | datasets, h, subjects |
| Model | FACED | KaggleERN | SEED-VIG | Sleep-EDF | CHB-MIT |
| BIOT | |||||
| CBraMod | |||||
| LaBraM-Base | |||||
| EEGPT-Large | |||||
| ST-EEGFormer-S | |||||
| REVE-Base |
| Dataset | Determined by | Head size | Spread | |
| FACED | dim/s at s | – M | ||
| KaggleERN | dim/s at s | – M | ||
| SEED-VIG | carried over from FACED | – M | ||
| Sleep-EDF | no-clamp bound | – M | ||
| CHB-MIT (15 Ch) | no-clamp bound | – M | ||
| CHB-MIT (2 Ch) | no-clamp bound | – M |
| Activation | at | Seeds escaping | Balanced accuracy |
| GELU | (and for ) | ||
| ReLU | |||
| ELU | |||
| LeakyReLU |
| Model | W | N1 | N2 | N3 | REM |
| BIOT † | |||||
| CBraMod | |||||
| LaBraM | |||||
| EEGPT | |||||
| ST-EEGFormer-S | |||||
| REVE-Base |
| Metric | Fold 1 | Fold 2 | Fold 3 | Fold 4 | Fold 5 | Mean | Positive | |
| Macro-F1 | ||||||||
| Accuracy | ||||||||
| Cohen’s |
| Model | Sleep-EDF (2 ch) | CHB-MIT (15 ch) | ||||||||||
| 30 s | 300 s | 600 s | 1,800 s | 3,600 s | GB | 30 s | 300 s | 600 s | 1,800 s | 3,600 s | GB | |
| EEGPT-Large | ||||||||||||
| ST-EEGFormer-S | – ∗ | |||||||||||
| BIOT | ||||||||||||
| LaBraM | ||||||||||||
| CBraMod | ||||||||||||
| Component | Shape | Values | KB (fp32) |
| Fast stream | 61,440 | 240.0 | |
| Slow stream | 21,888 | 85.5 | |
| Memory bank | 1,536 | 6.0 | |
| Writer queue | 4,096 | 16.0 | |
| Last output | 448 | 1.8 | |
| Total per channel | 89,408 | 349.2 |
| Sleep-EDF montage, two channels | CHB-MIT montage, fifteen channels | ||||||||
| History | ms | real time | GB peak | MB state | ms | real time | GB peak | MB state | |
| s | |||||||||
| min | |||||||||
| min | |||||||||
| min | |||||||||
| min | |||||||||