Spatiotemporal Hyperedges for EEG Seizure Detection and Prediction
Authors: Hyunju Kim, Sheo Yon Jhin, Noseong Park, Nabil Imam
Organizations: School of Computational Science and Engineering Georgia Institute of Technology Atlanta, GA, USA · School of Computing Korea Advanced Institute of Science and Technology Daejeon, Republic of Korea
Seizure detection and prediction from EEG are clinically important but challenging because seizures are rare, temporally localized, and propagate as coordinated events across multiple channels. Recent dynamic graph neural networks model this by running a temporal model over a sequence of per-time-step pairwise channel edges. However, this pairwise construction misses the spatiotemporal coupling that constitutes a seizure, at substantial training cost. We propose HyBrain, which summarizes spatiotemporal EEG evidence through a small set of soft hyperedges rather than pairwise edges. A per-channel Mamba backbone produces one token per (channel, second), and a spatiotemporal hyperedge block pools these tokens into E_h shared group embeddings through soft memberships and broadcasts them back. The same encoder serves three downstream tasks: window-based detection, one-second point-wise detection, and preictal seizure prediction. On TUSZ and CHB-MIT, HyBrain achieves the best AUROC on every reported setting against ten baselines, with the largest gap on long-clip preictal prediction. It also matches the most efficient baselines in training time and peak GPU memory. A qualitative analysis shows that even a single learned hyperedge cleanly captures the preictal -> ictal -> postictal trajectory on a real seizure clip.
Figures & tables
Figure 1. Membership of one spatiotemporal hyperedge over the channel-time tokens of a 60s TUSZ clip. The membership regime shifts across phases (dashed lines), coupling the tokens of each phase into a distinguishable group-level embedding. This emerges because our model learns hyperedges well-suited for seizure detection. Heatmap of one learned hyperedge's membership over the channel-time tokens of a 60-second TUSZ clip; the membership pattern is bursty in the preictal interval, flat near the mean during the annotated seizure, and uniformly low in the postictal interval.
Figure 2. Overview of HyBrain. A multi-channel EEG clip is first converted into one-second channel-time spectral tokens using a per-channel short-time Fourier transform (STFT). The channel-time token encoder then processes each channel independently with a Mamba backbone and adds a learnable channel embedding to preserve channel identity. The spatiotemporal hyperedge aggregation block pools all channel-time tokens into Eh learnable hyperedge embeddings through soft memberships and broadcasts the resulting group-level context back to the tokens, capturing high-order cross-channel and cross-time interactions. A temporal attention layer further refines intra-channel temporal dependencies, producing the shared representation G . Finally, task-specific heads adapt this shared representation to window-based seizure detection, point-wise seizure detection, and preictal seizure prediction. Block diagram of the \methodencoder showing per-channel STFT features, a Mamba-based channel-time token encoder with learnable channel embeddings, spatiotemporal hyperedge aggregation with soft memberships and hyperedge embeddings, temporal attention refinement, and task-specific heads for window-based detection, point-wise detection, and seizure prediction.
Method
Compute
Activation memory
GRU-style edge stream ( Gao and Ribeiro, 2022 )
O(TEd2)
O(TEd)
Mamba-style edge stream ( Kotoge et al., 2025 )
O(TEd)
O(TEd)
HyBrain hyperedge block
O(TNEhd)
O(TNEh+Ehd)
Table 1. Per-sample asymptotic cost of representative spatiotemporal interaction blocks.
Method
Time (s/ep.)
Memory (MB)
TUSZ (12 s)
TUSZ (60 s)
CHB-MIT (12 s)
AUROC
F1
AUROC
F1
AUROC
F1
LSTM
31.30
88.0
0.839 ± 0.016
0.403 ± 0.065
0.826 ± 0.050
0.427 ± 0.067
0.823 ± 0.075
0.064 ± 0.007
CNN-LSTM
29.90
1208.1
0.824 ± 0.014
0.365 ± 0.053
0.690 ± 0.026
0.242 ± 0.020
0.798 ± 0.110
0.049 ± 0.032
BIOT
86.07
1850.7
0.811 ± 0.006
0.301 ± 0.013
0.717 ± 0.060
0.292 ± 0.036
0.903 ± 0.018
0.054 ± 0.090
LaBraM
212.9
3576.7
0.863 ± 0.010
0.398 ± 0.034
0.865 ± 0.007
0.466 ± 0.016
0.787 ± 0.010
0.054 ± 0.022
EEGPT
151.2
1022.1
0.884 ± 0.004
0.438 ± 0.010
0.784 ± 0.009
0.343 ± 0.007
0.911 ± 0.016
0.266 ± 0.025
Table 2. Window-based seizure detection performance on TUSZ and CHB-MIT. Mean ± std across three seeds. The best result is highlighted in bold , and the second-best is underlined . Time denotes wall-clock seconds per training epoch on a single A6000, and Memory denotes peak GPU memory during training, reported in MB.
Method
Time (s/ep.)
Memory (MB)
TUSZ (12 s)
TUSZ (60 s)
CHB-MIT (12 s)
AUROC
F1
AUROC
F1
AUROC
F1
Dense-LSTM
28.5
415
0.869 ± 0.007
0.369 ± 0.031
0.863 ± 0.012
0.388 ± 0.028
0.857 ± 0.010
0.069 ± 0.016
Dense-CNN-LSTM
29.5
1804
0.873 ± 0.013
0.378 ± 0.044
0.887 ± 0.005
0.410 ± 0.002
0.914 ± 0.010
0.213 ± 0.026
Dense-BIOT
111.5
3889
0.858 ± 0.013
0.348 ± 0.027
0.834 ± 0.036
0.315 ± 0.041
0.841 ± 0.007
0.108 ± 0.016
Dense-GRU-GCN
30.0
5419
0.910 ± 0.003
0.517 ± 0.016
0.918 ± 0.007
0.554 ± 0.046
0.859 ± 0.003
0.086 ± 0.014
Dense-DCRNN
937.0
886
0.910 ± 0.004
0.499 ± 0.017
0.922 ± 0.004
0.486 ± 0.030
0.886 ± 0.020
0.148 ± 0.047
Table 3. Point-wise (per-second) seizure detection on TUSZ (12 s, 60 s) and CHB-MIT (12 s). All baselines are paper architectures adapted with a per-timestep readout head, trained with dense BCE on per-second labels.
Method
12 s
60 s
AUROC
F1
AUROC
F1
LSTM
0.488 ± 0.003
0.250 ± 0.019
0.659 ± 0.082
0.295 ± 0.019
CNN-LSTM
0.418 ± 0.009
0.244 ± 0.003
0.641 ± 0.019
0.357 ± 0.013
BIOT
0.497 ± 0.010
0.260 ± 0.027
0.613 ± 0.066
0.343 ± 0.057
LaBraM
0.464 ± 0.008
0.272 ± 0.010
0.509 ± 0.050
0.255 ± 0.045
EEGPT
0.457 ± 0.038
0.250 ± 0.062
0.640 ± 0.063
0.315 ± 0.056
Table 4. Seizure prediction (preictal vs. interictal binary classification) on TUSZ.
Method
TUSZ 12 s
TUSZ 60 s
AUROC
F1
AUROC
F1
HyBrain ( Eh=1 )
0.899 ± 0.006
0.551 ± 0.008
0.885 ± 0.010
0.564 ± 0.041
HyBrain ( Eh=2 )
0.894 ± 0.005
0.541 ± 0.025
0.892 ± 0.002
0.537 ± 0.041
HyBrain ( Eh=3 )
0.880 ± 0.006
0.540 ± 0.002
0.908 ± 0.006
0.636 ± 0.023
w/o hyperblock
0.888 ± 0.005
0.500 ± 0.028
0.866 ± 0.016
0.544 ± 0.038
Δ (vs. best)
−1.1 pt
−5.1 pt
−4.2 pt
−9.2 pt
Table 5. Hyperedge block ablation on TUSZ window-based detection. Δ rows are changes versus the best HyBrain configuration per column.
Method
Time (s/ep.)
Infer. (ms/seg.)
Memory (MB)
Memory cost
LSTM
31.30
0.013
88.0
0.26×
CNN-LSTM
29.90
0.264
1208.1
3.63×
BIOT
86.07
3.264
1850.7
5.56×
LaBraM
212.9
9.429
3576.7
10.75×
EEGPT
151.2
4.425
1022.1
3.07×
EvolveGCN
72.89
0.691
62.3
0.19×
Table 6. Computational efficiency of seizure detection on TUSZ 60 s. Inference denotes average milliseconds per EEG segment. Memory-cost values above 1 indicate higher cost than ours, whereas values below 1 indicate lower cost.
Figure 3. Accuracy-efficiency plots on TUSZ 60 s detection, showing AUROC versus training time per epoch (left) and peak GPU memory (right). Two scatter plots showing AUROC versus time (left) and AUROC versus GPU memory (right) for every baseline and \method. \methodoccupies the top-left region of both panels.
Figure 4. Membership of one learned hyperedge on a TUSZ 60 s clip that traverses the full preictal → ictal → postictal trajectory. (a) Per-channel topomaps for each phase. (b) Membership over time. Dashed lines mark the labeled ictal interval. A topomap and a heatmap showing hyperedge membership rising before the labeled seizure onset, peaking during the labeled ictal interval, and decaying after offset.
Figure 5. EvoBrain’s pairwise EEG connectivity visualization across interictal, preictal, and ictal states, shown alongside Figure 4 for comparison. Three side-by-side topographic plots reproducing EvoBrain's pairwise EEG connectivity visualization for interictal, preictal, and ictal states, drawn as channel-to-channel edges between electrode pairs.
Figure 6. Sensitivity of HyBrain to the number of hyperedges Eh on TUSZ 60 s detection. Left: AUROC. Right: F1. Two line plots showing AUROC and F1 as the number of hyperedges Eh changes from 1 to 9. The curves remain stable across the sweep and stay close to or above the baseline reference lines.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 7. Additional learned-hyperedge visualizations on TUSZ 60 s test clips. Rows (a)-(b) and (c)-(d) show preictal → ictal transitions; rows (e)-(f), (g)-(h), and (i)-(j) show ictal → postictal transitions. In every preictal → ictal example, the membership rises sharply across channels before or at the annotated onset; in every ictal → postictal example, the membership decays through the annotated offset. Five rows of paired topomap and heatmap visualizations showing the learned hyperedge membership during preictal-to-ictal and ictal-to-postictal transitions across different test clips.
Seizure diagnosis from EEG signals is a critical yet persistently challenging task, due to the complicated neural dynamics and the spurious connections in inter-channel modeling. While spatial-temporal graph neural networks (STGNNs) have advanced EEG brain network representation learning, the resulting graph structures suffer from low clinical plausibility and limited interpretability due to their purely data-driven nature. To this end, we introduce NeuroGRIP, a retrieval-augmented graph refinement framework that incorporates external medical knowledge to calibrate noisy EEG graphs. We first construct a large-scale, domain-specific knowledge base derived from authoritative clinical guidelines. Leveraging large language models, we extract structured biomedical entities and relations to form a textual knowledge graph (KG), which serves as external knowledge source of clinical priors. Our framework performs alignment-aware query construction by projecting STGNN-generated EEG node embeddings into the semantic space of KG. Semantic queries are then executed via FAISS-based similarity search over knowledge triplets to retrieve relation evidence. Each predicted edge is assigned a confidence score based on retrieved similarity, relation type, and source reliability, enabling us to prune medically implausible edges from the originally predicted graph. Extensive experiments on TUSZ and CHB-MIT demonstrate that NeuroGRIP not only improves seizure detection accuracy but also enhances interpretability by grounding each prediction in clinically validated knowledge. This work provides the first unified framework that tightly couples brain dynamics with external medical expertise via retrieval-augmented reasoning, paving the way for knowledge-enhanced, explainable clinical diagnosis. The code is available at: https://github.com/LincanLi-X/NeuroGRIP.
Lincan Li, Zheng Chen, Yushun Dong
Department of Computer Science, Florida State University, Tallahassee, United States · SANKEN, The University of Osaka, Japan.
Deep learning for EEG-based seizure detection faces critical challenges: severe annotation scarcity and extreme class imbalance, where ictal events comprise less than 10% of clinical recordings. We present DiffEEG, a 9.6M-parameter self-supervised foundation model that addresses both limitations through denoising diffusion pre-training and reinforcement learning (RL)-based fine-tuning. Pre-trained on 1.3M unlabeled segments from the Temple University Hospital Seizure Corpus (TUHSZ), DiffEEG learns generic neural representations via a 1D U-Net with multi-head self-attention. For downstream adaptation, a reinforced decision layer employs policy gradient optimization to directly maximize F1-score, prioritizing sensitivity to rare seizure events over overall accuracy. Under strict patient-wise evaluation (279 patients, Leave-One-Fold-Out), DiffEEG achieves 61% accuracy and 59% F1 for 4-class seizure subtyping, and 81% accuracy with 85% weighted F1 for binary detection, maintaining clinically viable seizure recall (59%) despite extreme imbalance (6.7% prevalence). Segment-level evaluation establishes an upper bound of 97.6% accuracy, confirming strong architectural capacity. DiffEEG demonstrates that diffusion-based pre-training combined with metric-aware reinforcement learning enables clinically deployable seizure monitoring with minimal labeled data requirements.
Abdulkader Helwan, Lina Abou-Abbas, Hussein El Amouri +2
Department of Electrical and Computer Engineering, Lebanese American University, Byblos, Lebanon · Institute of Applied Artificial Intelligence, T´ELUQ University, Montreal, Canada
Epilepsy is one of the most common neurological disorders globally, characterized by recurring seizures and significantly impacting the quality of life. Despite advancements in diagnostic techniques, the mitigation of risks faced by epilepsy patients remains challenging due to the unpredictability of seizure events. An accurate forecast of seizure onset helps to reduce risks in epilepsy patients. In this paper, we propose EEG-FuseFormer, a transformer-based feature fusion framework for seizure-onset prediction that combines intermediate features extracted from Convolutional Neural Networks-Long Short-Term Memory (CNN-LSTM) and ResNet-18 networks. The CNN-LSTM architecture captures both spatial and temporal features directly from the raw signal, whereas the ResNet-18 extracts features from the Short-Time Fourier Transform (STFT) representation of the EEG signals. Fusion is carried out using a transformer encoder, and the final prediction is generated using fully connected dense layers. The CHB-MIT dataset was used to validate the proposed model. The results show that the proposed model achieves a mean recall of 98.85% and outperforms most of the state-of-the-art methods. This study evaluates the ability of the proposed feature fusion model to generalize in cross-patient testing scenarios. Fine-tuning pre-trained models on limited target patient data (target adaptation) within the cross-patient validation framework results in higher recall, precision, and F1-score metrics in comparison to the conventional cross-patient validation approach. Finally, the runtime-based computational complexity of the model is assessed across diverse hardware platforms to highlight the performance-complexity trade-off.
Vigneshwar Hariharan, Chithra Reghuvaran, Arlene John +4
National University of Singapore · University College Dublin · University of Twente +1