What masking geometry works best for EEG foundation models?
Organizations: Donders Institute for Brain, Cognition and Behaviour Radboud University Nijmegen, The Netherlands · Yneuro, University of California San Diego, Paris, France · Lab-STICC, IMT Atlantique Brest, France · SCCN, INC, SDSC University of California San Diego, USA CNRS, France · Université Paris-Saclay, Inria, CEA Palaiseau, France
Abstract
EEG foundation models hold promise for scalable brain-signal decoding across clinical and cognitive neuroscience applications, yet their pre-training pipelines remain poorly understood. Among design choices, the masking strategy is particularly critical: it determines what the network must predict and from which context. Yet it has never been ablated in isolation, as each new model bundles a new masking strategy with a new backbone and objective. In this paper, we formalize the design choices for spatio-temporal masking strategies and train various models with a single pipeline under varying masking configurations across two SSL frameworks (MAE and JEPA). We then systematically evaluate the resulting 58 pre-trained models on the 12 datasets of OpenEEGBench under a linear probe. Both frameworks agree on an optimal masking configuration and on shared failure modes. Outside these, performance is robust: 11 MAE and 9 JEPA configurations are statistically indistinguishable from the best. We further identify a novel JEPA-specific failure mode, tagged bias-inflation collapse, invisible to standard detectors. With a well-chosen mask, our pipeline reaches REVE-level downstream performance at a fraction of REVE's pre-training compute.
Figures & tables
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | # params |
|---|---|
| Linear patch embedding ( , with bias) | 102,912 |
| Encoder layer (single, ): | 3,146,240 |
| Self-attention , | 1,048,576 |
| GEGLU FFN (gate + value, output projections) | 2,096,640 |
| RMSNorm scales | 1,024 |
| Encoder transformer total (4 layers) | 12,584,960 |
| Hyperparameter | Value |
|---|---|
| Optimizer | AdamW |
| (PyTorch default) | |
| weight decay | |
| Peak learning rate | |
| Final learning rate | |
| Warmup schedule | linear, factor |
| Framework | Mean | Median | SD | Min–Max | |
|---|---|---|---|---|---|
| JEPA | – | ||||
| MAE | – | ||||
| Overall | – |
| Temporal block length (patches) | ||||||
| Spatial radius | ||||||
| ( "one" ) | 843 | 422 | 211 | 106 | 53 | 26 |
| 220 | 110 | 55 | 28 | 14 | 7 | |
| 98 | 49 | 24 | 12 | 6 | 3 | |
| 55 | 28 | 14 | 7 | 4 | 2 | |
| ( "all" ) | 26 | 13 | 7 | 3 | 2 | |
| Temporal block length (patches) | |||||||
|---|---|---|---|---|---|---|---|
| Spatial radius | metric | ||||||
| ( "one" ) | ch / sphere | ||||||
| tot % | |||||||
| ch / sphere | |||||||
| tot % | |||||||
| ch / sphere | |||||||
| Dataset | classes | Distribution (%) | |
|---|---|---|---|
| arithmetic_zyma2019 | |||
| bcic2020-3 | |||
| bcic2a | |||
| chbmit | |||
| faced | |||
| isruc-sleep |
| Dataset | projected? | ||||
|---|---|---|---|---|---|
| arithmetic_zyma2019 | yes | ||||
| bcic2a | yes | ||||
| bcic2020-3 | yes | ||||
| physionet | yes | ||||
| chbmit | yes | ||||
| faced | yes |
| min | max | |
|---|---|---|
| Dataset | ||
| arithmetic_zyma2019 | 0.595 | 0.778 |
| bcic2020-3 | 0.231 | 0.294 |
| bcic2a | 0.263 | 0.475 |
| chbmit | 0.771 | 0.922 |
| faced | 0.158 | 0.332 |
| Split | mean | std | range | |
|---|---|---|---|---|
| train | ||||
| val | ||||
| test |
| Predictor | test |
|---|---|
| DummyRegressor (predicts ) | |
| REVE baseline, ridge probe (mean over 5 seeds) | |
| Gap above the dummy floor |
| Band / dataset | MAE | JEPA | JEPA , collapsed | Untrained |
|---|---|---|---|---|
| band | ||||
| arithmetic_zyma2019 | ||||
| bcic2020-3 | ||||
| bcic2a | ||||
| chbmit | ||||
| faced | ||||
| Ours , ridge | Public, ridge | Public, SGD (published) | |||||
|---|---|---|---|---|---|---|---|
| Dataset | MAE | JEPA | REVE-Base | CBraMod | EEGPT | LaBraM | BIOT |
| arithmetic_zyma2019 | |||||||
| bcic2020-3 | |||||||
| bcic2a | |||||||
| chbmit | — | — | — | ||||
| faced | |||||||
| Backbone | Published mask | Recommended geometry | Recommended better on |
|---|---|---|---|
| REVE-Small (core grid) | of | ||
| CBraMod | of | ||
| LaBraM | of |
| Framework | Geometry | (grid) | ||||
|---|---|---|---|---|---|---|
| MAE | ||||||
| MAE | ||||||
| MAE | — | |||||
| MAE | ||||||
| JEPA | ||||||
| JEPA |
| Run | Patch | Framework | Token dim. | Mask | , ep. | Params enc. / tot. | GPU-h | OEB score | ||
|---|---|---|---|---|---|---|---|---|---|---|
| (A) same model size | s | JEPA | / | / | / | M / M | ||||
| (B) same features per second | s | JEPA | / | / | / | M / M | ||||
| Core grid, collapsed | s | JEPA | / | / | / | M / M | ||||
| Core grid, collapsed | s | JEPA | / | / | / | M / M | ||||
| Core grid, healthy | s | JEPA | / | / | / | M / M | ||||
| Core grid, MAE control | s | MAE | / | / | / | M / M |