AutoBCI: Forecast-Guided Agentic Neural Architecture Discovery for EEG-Based Brain--Computer Interfaces
Organizations: Nanyang Technological University, Singapore · The University of Hong Kong, Hong Kong, China · Southeast University, China
Abstract
EEG-based brain-computer interfaces support a broad range of applications, yet designing decoding architectures that perform well across diverse tasks remains challenging. We introduce AutoBCI, an agentic framework in which a Designer Agent and a Forecaster Agent support the discovery and selection of EEG decoding architectures across tasks. The Designer Agent performs Pool-Guided Architecture Discovery (PGAD), generating and refining architectures through training and validation across multiple EEG tasks, such as emotion recognition, motor imagery, and sleep staging. The Forecaster Agent performs Performance Estimation from Early Knowledge (PEEK), using architecture code, the training protocol, and early learning curves to predict full-budget validation performance and select promising candidates for continued training. Across 14 EEG datasets spanning motor imagery, emotion recognition, and sleep staging, we evaluate AutoBCI with six LLMs, including Opus 5.5 and GPT 5.6 Sol, and compare the architectures selected by the search procedure against ten baselines: six conventional EEG models and four foundation models. The architecture discovered by AutoBCI with Claude Opus 5.5 achieves 64.16% average test balanced accuracy (bAcc), compared with 63.87% for REVE, the strongest baseline on this metric. Using ten observed epochs, PEEK reduces mean absolute error in predicting average validation bAcc from 2.20 to 1.36 percentage points, a 38.1% reduction relative to the best-observed-score baseline.
Figures & tables
| Task | No. | Datasets | Unified input | Classes |
|---|---|---|---|---|
| Motor Imagery | 1 2 3 4 5 6 7 8 | BCIC-IV-2a ( Tangermann et al., 2012 ) OpenBMI-MI ( Lee et al., 2019 ) BCIC-Upperlimb ( Jeong et al., 2022 ) Cho2017 ( Cho et al., 2017 ) HighGamma ( Schirrmeister et al., 2017 ) PhysioNet-MI ( Schalk et al., 2004 ) SHU-MI ( Ma et al., 2022 ) Shin2017A ( Shin et al., 2016 ) | 65 ch 4 s | Left, Right, Foot, Tongue, Cylin, Sphe, Lumbrical (7 classes) |
| Emotion Recognition | 9 10 11 | SEED ( Duan et al., 2013 ) SEED-IV ( Zheng et al., 2018 ) SEED-V ( Liu et al., 2021 ) | 65 ch 4 s | Positive, Neutral, Negative, Sad, Fear, Happy, Disgust (7 classes) |
| Sleep Staging | 12 13 14 | Sleep-EDF ( Kemp et al., 2000 ) ISRUC ( Khalighi et al., 2016 ) HMC ( Alvarez-Estevez and Rijsman, 2021 ) | 8 ch 30 s | Wake, N1, N2, N3, REM (5 classes) |
| Model tier | Provider | LLM | Agent interface | Input (USD/1M) | Output (USD/1M) |
| Flash | Anthropic | Claude Sonnet 5 | Claude Code | 2.00 | 10.00 |
| DeepSeek | DeepSeek V4.1 Flash | Claude Code | 0.30 | 1.20 | |
| Gemini 3.8 Flash | Gemini CLI | 0.75 | 3.75 | ||
| Alibaba | Qwen3.8 Flash | Claude Code | 0.15 | 0.47 | |
| Flagship | OpenAI | GPT-5.6 Sol | Codex CLI | 4.00 | 20.00 |
| Anthropic | Claude Opus 5.5 | Claude Code | 4.00 | 20.00 |
| Model | Spec | Motor Imagery | Emotion | Sleep | Average | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| bAcc | wF1 | Kappa | bAcc | wF1 | Kappa | bAcc | wF1 | Kappa | bAcc | wF1 | Kappa | ||
| EEGNet ( Lawhern et al., 2018 ) | 8.8K | 57.83 | 57.12 | 0.233 | 37.97 | 34.48 | 0.164 | 69.09 | 70.60 | 0.640 | 54.96 | 54.07 | 0.345 |
| TSception ( Ding et al., 2022 ) | 17.5K | 54.58 | 53.63 | 0.167 | 43.93 | 45.91 | 0.256 | 66.75 | 70.74 | 0.648 | 55.09 | 56.76 | 0.357 |
| DCN ( Schirrmeister et al., 2017 ) | 317K | 66.70 | 66.36 | 0.396 | 42.73 | 38.03 | 0.215 | 73.89 | 76.68 | 0.697 | 61.10 | 60.35 | 0.436 |
| FAST ( Jiang et al., 2026 ) | 475K | 58.59 | 58.31 | 0.243 | 35.39 | 36.94 | 0.137 | 72.26 | 74.99 | 0.686 | 55.41 | 56.75 | 0.355 |
| EEG-Conformer ( Song et al., 2022 ) | 2.13M | 66.32 | 66.21 | 0.396 | 48.08 | 48.92 | 0.302 | 72.34 | 76.07 | 0.701 | 62.25 | 63.73 | 0.466 |
| Model | Round-best declines | Older-parent rounds | Older-parent link in champion lineage |
|---|---|---|---|
| Claude Opus 5.5 | None | R3 | Yes ( ) |
| Claude Sonnet 5 | R4 (0.221 pp) | R3, R4, R5, R6 | Yes ( ) |
| DeepSeek V4.1 Flash | R2 (0.030 pp) | R3, R4, R6 | Yes ( ) |
| Gemini 3.8 Flash | None | R3 | None (fresh) |
| GPT-5.6 Sol | None | R3 | Yes ( ) |
| Qwen3.8 Flash | None | R5 | None (fresh) |
| Opus candidates | Sonnet candidates | |||
|---|---|---|---|---|
| Forecaster | Top-1 | Top-2 | Top-1 | Top-2 |
| Best observed score | 50.0 | 83.3 | 50.0 | 55.6 |
| Self | 50.0 | 77.8 | 50.0 | 72.2 |
| DeepSeek V4.1 Flash | 50.0 | 77.8 | 50.0 | 66.7 |
| GPT-5.6 Sol | 44.4 | 77.8 | 50.0 | 66.7 |
| Gemini 3.8 Flash | 44.4 | 83.3 | 55.6 | 61.1 |
| Method | Winner retention | Full runs | Task-epochs | Epoch savings | GPU job-hours |
|---|---|---|---|---|---|
| Full training | 100.0 | 807 | 32,280 | 0.0 | 68.58 |
| PEEK ( ) | 61.1 | 108 | 11,310 | 65.0 | 24.52 |
| PEEK ( ) | 83.3 | 216 | 14,550 | 54.9 | 31.86 |
| PEEK ( ) | 91.7 | 324 | 17,790 | 44.9 | 38.85 |
| PEEK ( ) | 94.4 | 429 | 20,940 | 35.1 | 45.53 |
| PEEK ( ) | 97.2 | 534 | 24,090 | 25.4 | 52.18 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Dataset | Train | Validation | Test |
|---|---|---|---|---|
| Emotion | SEED | Trials 1–9 | Trials 10–12 | Trials 13–15 |
| SEED-IV | Trials 1–16 | Trials 17–20 | Trials 21–24 | |
| SEED-V | Trials 1–5 | Trials 6–10 | Trials 11–15 | |
| MI | BCI Competition IV-2a | Subjects 0–6 | Subjects 0–6 | Subjects 7–8 |
| BCI Upper Limb | Subjects 0–10 | Subjects 0–10 | Subjects 11–14 | |
| Cho2017 | Subjects 0–39 | Subjects 0–39 | Subjects 40–48 |
| Task | Dataset | Classes | Train | Validation | Test |
|---|---|---|---|---|---|
| Emotion | SEED (3 classes) | 3 | 22,545 | 7,905 | 7,620 |
| SEED-IV | 4 | 26,025 | 5,340 | 6,210 | |
| SEED-V | 5 | 8,512 | 10,672 | 9,984 | |
| MI | BCI Competition IV-2a | 4 | 3,148 | 788 | 1,152 |
| BCI Upper Limb | 3 | 2,640 | 660 | 1,200 | |
| Cho2017 | 2 | 6,464 | 1,616 | 1,800 |
| Setting | Value |
|---|---|
| Full training budget | 40 epochs per task |
| Batch size | 512 |
| Optimizer | Fused AdamW |
| Initial learning rate | |
| Weight decay | |
| Adam coefficients |
| Model | Precision | GFLOPs | Time (s) | MiB | Test bAcc (%) |
|---|---|---|---|---|---|
| EEGNet | FP32 | 0.126 | 644 | ||
| BF16 mixed | 0.126 | 371 | |||
| BF16 full | 0.126 | 345 | |||
| DeepConvNet | FP32 | 0.328 | 1336 | ||
| BF16 mixed | 0.328 | 1099 | |||
| BF16 full | 0.328 | 1083 |