Decoupling Time and Space: A Temporally Conditioned Refinement for EEG Source Imaging
Authors: Marco Morik, Jesse Palarus, Carmen Vidaurre, Klaus-Robert Müller, Shinichi Nakajima
Organizations: Berlin Institute for the Foundations of Learning and Data, Germany · Machine Learning Group, Technische Universität Berlin, Germany · Basque Center on Cognition, Brain and Language, Mikeletegi Pasealekua, Spain · Ikerbasque, Bizkaia, Spain · Department of Artificial Intelligence, Korea University, Seoul, Korea · Max Planck Institut für Informatik, Saarbrücken, Germany · RIKEN AIP, Tokyo, Japan
Electroencephalography (EEG) offers millisecond temporal resolution, but inferring underlying neural sources is a severely ill-posed spatial inverse problem. While deep learning has advanced spatial reconstruction, current architectures face a critical dilemma: frame-by-frame models discard vital temporal context, whereas full 4D spatiotemporal networks introduce an architectural trade-off between reconstruction accuracy and inference cost. We propose a novel two-stream framework that explicitly decouples global temporal representation learning from per-time-point spatial refinement. A Transformer-based Temporal Condition Encoder processes the entire EEG sequence via factorized spatiotemporal attention, retaining sensor-resolved features. A fixed inverse then maps these features into source-indexed conditioning for a per-timestep Source-Space Transformer or volumetric convolutional refiner. Extensive evaluations on realistic synthetic data demonstrate that this temporal prior dramatically improves spatial localization, outperforming classical and spatiotemporal baselines, particularly in high-noise and multi-source regimes. Training across diverse leadfields and explicit operator mismatches improves transfer to unseen head geometries and brings template-based reconstruction closer to subject-specific inversion. Furthermore, we apply the model trained only on synthetic EEG data to real-world EEG. A logistic regressor fit on source power differences in eyes-open, eyes-closed conditions successfully decodes age groups.
Figures & tables
Fig. 1 : Temporal conditioning for source refinement. The novel Temporal Conditioning (+TC) path (highlighted in blue) applies the Temporal Condition Encoder gθ to the full EEG sequence, alternates temporal and sensor attention, and unpatches features to an aligned sequence of condition vectors c . At each time point, the subject-specific eLORETA-inverse Lj† transports the EEG slice and each condition feature into source space. The refiner fθ fuses these two source-space inputs to predict the final sources x^t .
Space
Head model
M
S
# (train/eval)
Surface
LEMON [ 53 ]
61
516
127/10
IBP2021 [ 54 ]
62
516
11/5
fsaverage [ 55 ]
16-62
516
2-3/1
Volume
fsaverage [ 55 ]
16-62
599
3/1
IBP2021 [ 54 ]
62
461-641
11/5
TABLE I : Lead-field pools. M / S denote sensor/source-location counts, the final column counts training/evaluation lead fields. Surface and volume spaces use oct4 and 14-mm grids, respectively.
Fig. 2 : Reconstructed magnitudes at one active patch center on an unseen IBP2021 head at 5dB SNR. SoST+TC (ours), SoST, DeepSIF, and eLORETA process the same sequence. The displayed 0-500 ms window is taken from the 512-sample, 256 Hz record.
Method
NMAE
NEMD
Spatial NEMD
Temporal NEMD
Cosine Error
Params. (M)
Latency (ms)
Surface source space
SoST
0.824
1.023
0.804
0.572
0.354
16.5
489.8
SoST+TC
0.645
0.701
0.652
0.255
0.219
17.4
504.8
SoST+TC+OM
0.577
0.641
0.593
0.232
0.185
17.4
505.9
Deep CB [ 39 ]
0.889
0.844
0.819
0.404
0.380
<0.1
50.2
DeepSIF [ 19 ]
0.860
0.937
0.873
0.728
0.399
7.4
19.3
TABLE II : Aggregate reconstruction errors (lower is better) and computational cost. Evaluation uses 200 simulated sequences per head (512 samples at 256 Hz), with SNR, active patch count, and patch standard deviation sampled uniformly from [−10,30] dB, {1,…,20} , and [5,40] mm, respectively. Within each surface seed, errors are averaged over ten held-out LEMON and five held-out IBP2021 heads separately, then equally across datasets, taking the mean over three training seeds. Volume rows average five held-out IBP2021 heads for one training seed. Runtime protocol: FP32, batch size one, 512 time points per forward call, 50 timed runs on an NVIDIA A100 80GB.
Fig. 3 : Robustness to sensor noise and source multiplicity. Rows show NMAE, spatial NEMD, and temporal NEMD. The SNR sweep (left) draws source counts uniformly from 1-10, the source-count sweep (right) fixes SNR at 5 dB. Each point first averages 15 equally weighted held-out heads (ten LEMON, five IBP2021) within each seed, then reports the mean ± sample SD over three training seeds. Our proposed SoST+TC retains lower errors across both the evaluated noise and source-count settings.
Fig. 4 : Conditioning interventions on one held-out IBP2021 head, with 64 paired simulations per difficulty. Left: spatial NEMD, right: temporal NEMD. Positive deltas indicate worse errors than the intact input. Boxes show medians and interquartile range, and points show individual simulations. Interventions are zero initial source estimate input, zero condition, and time-shuffled condition. Easy/medium/hard use 15/5/ −5 dB SNR and 1/1-5/5-10 active sources.
Fig. 5 : Head-model generalization. A: NMAE (lower is better) averages ten held-out LEMON or five held-out IBP2021 heads in the respective columns, then reports seed mean ± sample SD over three seeds. Parentheses identify single-pool training. Specialists transfer less accurately across datasets, mixed training reduces transfer loss. Additionally, Operator Mixing improves the template inverse performance. B: Maps show full-epoch source-vector RMS divided by one common ground-truth maximum. The mismatched specialist underestimates activity and misses one of three active regions.
Fig. 6 : LEMON age-group decoding on 10 unseen test subjects, whose geometries were held out from source-model training. A: Descriptive older-minus-younger Cohen’s d maps from these ten participants, using pooled within-age SD within the 4-35 Hz frequency band. B: Test balanced accuracy of separate logistic classifiers trained on 115 participants using one indicated band or total-power reactivity (516 source features, 61 raw-EEG features where missing electrodes use training-fold mean imputation). SoST+TC shows great age-group decoding.
High-density electroencephalography (HD-EEG) enables fine-grained measurement of cortical activity but requires expensive hardware and lengthy setup times, limiting its clinical and research accessibility. We propose EMAG (EEG Mixture of Anisotropic Gaussians), a differentiable framework that reconstructs HD-EEG signals from a sparse subset of low-density (LD) electrodes by representing brain electrical sources as a mixture of anisotropic 4D space-time Gaussians. EMAG places a mixture of multiple Gaussians at each point of a spherical brain grid, each parameterized by a full 4 x 4 precision matrix, enabling anisotropic spatial spreads and explicit coupling between spatial and temporal dimensions. The forward model renders scalp EEG via differentiable Gaussian field contributions at electrode locations, enabling end-to-end training without explicit source localization supervision. We evaluate EMAG on three public EEG benchmarks (Localize-MI, SEED, and SEED-IV) at super-resolution factors of 2x through 8/16x. EMAG outperforms the current state-of-the-art EEG super-resolution method at most super-resolution factors on three standard benchmarks (Localize-MI, SEED, SEED-IV). The explicit Gaussian parameterization further enables direct visualization and interpretability of learned brain source configurations, potentially opening avenues for clinical and neuroscientific applications, such as source localization or biomarker discovery.
Alex Lazarovich, Ofir Itzhak Shahar, Gur Elkin +1
Stein Faculty of Computer and Information Science Ben-Gurion University of the Negev, Israel
Electroencephalography (EEG) is a widely adopted technique for monitoring brain activity, offering valuable insights into neurological states due to its high temporal resolution and cost-effectiveness. To enhance the analysis of complex EEG data, we propose EEG-TransNet, an architecture designed to capture temporal, regional, and synchronous features of EEG signals. EEG-TransNet introduces three key modules: 1) a preprocessing and feature extraction module leveraging ResNet and wavelet-based denoising, 2) a Local Self-Attention Block for regional feature learning, and 3) a Fuzzy-Attention Synchronous Transformer (FAST) to model spatiotemporal dependencies. Through extensive experiments on three EEG datasets (BETA, SEED, and DepEEG), the proposed model consistently outperforms other methods in terms of classification accuracy and robustness across varying signal lengths. Ablation studies confirm the contribution of the Local Self-Attention Block in improving performance, and the inclusion of depthwise separable convolutions in the decoder reduces computational complexity while maintaining high accuracy. EEG-TransNet's ability to generalize across subjects with minimal performance variation highlights its potential as a robust tool for EEG-based brain activity classification and emotion recognition tasks.
Xinglong Cui, Dian Gu
Beijing Neurodeep Technology Co., Ltd, Beijing, 100176, China · University of Pennsylvania, Seattle, 98121, USA
Electroencephalography (EEG) generation is essential for alleviating data scarcity and enabling large scale neural modeling in brain computer interface applications. However, existing flow based approaches assume that every channel and every time segment within a sample shares a single global time progression, overlooking the fact that not all EEG moments are equal. To address this overlooked heterogeneity, we propose an adaptive EEG generation framework built on conditional flow matching. The framework introduces Position-Adaptive Time Scheduling, which tracks per position reconstruction error to modulate a position specific time progress within the flow matching trajectory. It further incorporates Factorized Spatio-Temporal Attention and a frequency aligned multi resolution spectral consistency loss to model inter channel dependencies induced by volume conduction and compensate for the power law spectral bias of EEG, thereby improving the quality of generated signals. Extensive experiments on three EEG datasets with distinct acquisition protocols and task semantics show that our framework consistently outperforms the strongest baseline, reducing TS-FID by up to 62.2% and improving downstream classification accuracy gain by up to 6.77 percentage points. These results suggest that the proposed method represents a promising step toward scalable, high fidelity data augmentation for real world brain computer interface applications.
Boheng Liu, Ziyu Li, Chenghua Duan +2
School of Computer Science and Technology Beijing Institute of Technology Beijing, China