RPA: Residual Patch-Token Adapter for Image Retrieval from EEG and MEG
Organizations: Vortek Lab Inc. · Department of Computer Science and Technology, Tsinghua University · Biology and Biological Engineering, California Institute of Technology
Abstract
Most existing MEG and EEG (M/EEG) visual decoding methods align brain signals with a single global embedding extracted from a pretrained visual encoder, leaving open whether intermediate patch representations, which preserve richer and more granular rich visual information, can improve representation learning. To address this question, we introduce the Residual Patch Adapter (RPA), a lightweight, modular adapter that leverages all patch tokens from an intermediate layer of a ViT visual encoder for alignment. Through extensive ablation analyses, we first show that pooling or masking patch tokens degrades the learned representation, demonstrating that retaining the full set of patch tokens is important for EEG alignment, while the CLS token provides little unique information. We then use a series of six quantitative feature analyses to show that both higher-level semantics and lower-level visual features, including color and texture, are essential for this EEG-to-image alignment. Under current protocols, our system achieves Top-1 accuracies of 95.4% within-subject and 35.5% cross-subject on THINGS-EEG2, and 65.2% and 6.7%, respectively, on THINGS-MEG, achieving state-of-the-art (SOTA) performance across both datasets. Evaluations with alternative brain encoders, including pretrained EEG foundation models, demonstrate that the approach extends beyond the projection-based EEG encoder. Furthermore, we provide a plug-and-play interface that allows RPA to be replaced by convolution, attention, or ConvNeXt alternatives. Together, these findings provide significant insight into M/EEG-to-image representation learning by establishing design principles for leveraging the latent space of visual encoders, and open new directions for brain--image alignment and non-invasive brain--computer interface (BCI).
Figures & tables
| EEG | MEG | |||||||
| Within | Cross | Within | Cross | |||||
| Method | T1 | T5 | T1 | T5 | T1 | T5 | T1 | T5 |
| NICE ( Song et al., 2024 ) | 13.8 | 39.5 | 6.2 | 21.4 | 12.8 | 36.0 | – | – |
| ATM ( Li et al., 2024 ) | 28.6 | 58.5 | 11.8 | 33.7 | – | – | – | – |
| UBP ( Wu et al., 2025 ) | 50.9 | 79.7 | 12.4 | 33.4 | 26.7 | 55.2 | 2.2 | 10.4 |
| BrainHIVE ( Zheng et al., 2026 ) | 75.7 | 94.6 | 20.0 | 44.1 | 33.7 | 60.5 | 5.4 | 15.2 |
| Low-level | High-level | ||||||
|---|---|---|---|---|---|---|---|
| Method | PixCorr | SSIM | AlexNet-2 | AlexNet-5 | Inception | CLIP | SwAV |
| CognitionCapturer (All) ( Zhang et al., 2025 ) | 0.150 | 0.347 | 0.754 | 0.623 | 0.669 | 0.715 | 0.590 |
| AVDE ( Dai et al., 2026 ) | 0.147 | 0.366 | 0.766 | 0.835 | 0.724 | 0.747 | 0.587 |
| BrainHIVE ( Zheng et al., 2026 ) | 0.195 | 0.336 | 0.843 | 0.905 | 0.756 | 0.808 | 0.554 |
| Ours | 0.225 | 0.454 | 0.865 | 0.909 | 0.783 | 0.811 | 0.526 |
| Visual input | Visual adapter | Top-1 (%) |
|---|---|---|
| CLS token | LayerNorm + linear | |
| patch tokens | mean, LayerNorm + linear | |
| patch tokens | pointwise RPA | |
| patch tokens | local RPA |
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Within subject | LOSO |
|---|---|---|
| Training responses | 16,540 from one subject | 148,860 from nine subjects |
| Test subject | same subject, unseen concepts | held-out subject, unseen concepts |
| Default window | 0–700 ms | 0–1,000 ms |
| Epochs | 25 | 15 |
| Peak learning rate | ||
| Projection dropout | 0.5 | 0.7 |
| Method | S1 | S2 | S3 | S4 | S5 | S6 | S7 | S8 | S9 | S10 | Mean |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Top-1 accuracy (%) | |||||||||||
| NICE ( Song et al., 2024 ) | 12.3 | 10.4 | 13.1 | 16.4 | 8.0 | 14.1 | 15.2 | 20.0 | 13.3 | 14.9 | 13.8 |
| UBP ( Wu et al., 2025 ) | 41.2 | 51.2 | 51.2 | 51.1 | 42.2 | 57.5 | 49.0 | 58.6 | 45.1 | 61.5 | 50.9 |
| ViEEG ( Liu et al., 2026a ) | 34.1 | 38.4 | 40.6 | 50.1 | 28.9 | 44.3 | 38.6 | 54.0 | 37.3 | 42.8 | 40.9 |
| NeuroBridge ( Zhang et al., 2026 ) | 50.0 | 63.2 | 61.6 | 61.4 | 54.8 | 69.7 | 62.7 | 71.2 | 64.0 | 73.6 | 63.2 |
| BrainHIVE ( Zheng et al., 2026 ) | 64.3 | 76.3 | 74.0 | 67.0 | 68.0 | 81.5 | 76.8 | 84.8 | 76.8 | 87.3 | 75.7 |
| Method | S1 | S2 | S3 | S4 | S5 | S6 | S7 | S8 | S9 | S10 | Mean |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Top-1 accuracy (%) | |||||||||||
| NICE ( Song et al., 2024 ) | 7.6 | 5.9 | 6.0 | 6.3 | 4.4 | 5.6 | 5.6 | 6.3 | 5.7 | 8.4 | 6.2 |
| UBP ( Wu et al., 2025 ) | 11.5 | 15.5 | 9.8 | 13.0 | 8.8 | 11.7 | 10.2 | 12.2 | 15.5 | 16.0 | 12.4 |
| ViEEG ( Liu et al., 2026a ) | 22.7 | 24.7 | 19.0 | 25.5 | 19.8 | 20.7 | 20.9 | 20.8 | 23.8 | 31.2 | 22.9 |
| NeuroBridge ( Zhang et al., 2026 ) | 23.2 | 21.2 | 13.2 | 17.0 | 14.5 | 25.0 | 15.3 | 20.1 | 13.7 | 27.2 | 19.0 |
| BrainHIVE ( Zheng et al., 2026 ) | – | – | – | – | – | – | – | – | – | – | 20.0 |
| Method | S1 | S2 | S3 | S4 | Mean |
|---|---|---|---|---|---|
| Top-1 accuracy (%) | |||||
| NICE ( Song et al., 2024 ) | 9.6 | 18.5 | 14.2 | 9.0 | 12.8 |
| UBP ( Wu et al., 2025 ) | 15.0 | 46.0 | 27.3 | 18.5 | 26.7 |
| ViEEG ( Liu et al., 2026a ) | 16.6 | 37.4 | 30.6 | 17.2 | 25.5 |
| NeuroBridge ( Zhang et al., 2026 ) | 16.5 | 53.7 | 40.4 | 18.1 | 32.2 |
| BrainHIVE ( Zheng et al., 2026 ) | 14.0 | 63.8 | 41.0 | 17.0 | 33.9 |
| Method | S1 | S2 | S3 | S4 | Mean |
|---|---|---|---|---|---|
| Top-1 accuracy (%) | |||||
| UBP ( Wu et al., 2025 ) | 2.0 | 1.5 | 2.7 | 2.5 | 2.2 |
| NeuroBridge ( Zhang et al., 2026 ) | 4.3 | 3.6 | 3.0 | 2.5 | 3.4 |
| BrainHIVE ( Zheng et al., 2026 ) | – | – | – | – | 5.4 |
| Blur ( Liu et al., 2026b ) | 2.9 | 7.7 | 5.8 | 4.7 | 5.3 |
| Shallow Align. ( Du et al., 2026 ) | 1.3 | 6.6 | 5.4 | 1.5 | 3.7 |
| Family | Adapter | Within | LOSO |
|---|---|---|---|
| RPA | local | ||
| RPA | pointwise | ||
| RPA ablation | no final GELU | ||
| RPA ablation | pointwise + query pool | ||
| scalar attention | residual-MLP values | ||
| scalar attention | MLP values |
| Adapter | Parameters | Adapter | Parameters |
|---|---|---|---|
| CLS-token / patch-mean controls | 1.314M | patch mean + linear (no LN) | 1.312M |
| pointwise RPA | 2.038M | local RPA | 2.562M |
| local RPA, no final GELU | 2.562M | depthwise-separable conv | 2.969M |
| query pool + MLP values | 3.610M | pointwise RPA + query pool | 3.679M |
| query pool + linear values | 4.592M | 8-head MHA query pool | 7.872M |
| full-width ConvNeXt | 14.493M | transformer + MHA pool | 20.993M |
| Adapter | Adapter output before L2 normalization |
|---|---|
| local RPA | |
| pointwise RPA | |
| local RPA, no final GELU | local RPA with the last GELU in removed; the residual is retained |
| pointwise RPA + query pool | ; |
| query pool + residual-MLP values | , , , |
| query pool + MLP values | , with the same widths |
| Adapter | Structure | Params | Within subject | LOSO |
|---|---|---|---|---|
| width 64 | 1.52M | 93.4 | 34.9 | |
| no local mixing | 2.04M | 93.0 | 34.8 | |
| reference | 2.56M | 95.4 | 35.5 | |
| all-local | 7.81M | 91.9 | 33.8 | |
| width 1280 | 19.35M | 94.5 | 35.3 |
| Encoder | rank 1 | rank 2 | rank 3 | rank 4 | rank 5 | rank 6 (final) |
|---|---|---|---|---|---|---|
| EEG within-subject | ||||||
| CLIP-H | ||||||
| I-JEPA | ||||||
| DINOv2-L | ||||||
| DINOv2-S | ||||||
| ViT-B/32 | ||||||
| Encoder | Pretraining | Final layer | Test-best sampled layer (rank) |
|---|---|---|---|
| CLIP ViT-H/14 | language contrastive | (r3) | |
| I-JEPA ViT-H/14 | masked SSL | (r3) | |
| DINOv2-L/14 | SSL distillation | (r3) | |
| DINOv2-S/14 | SSL distillation | (r2) | |
| CLIP ViT-B/32 | language contrastive | (r2) | |
| CLIP RN50 | language contrastive | (r5) |
| Analysis | Vector or comparison | Result |
|---|---|---|
| Centered cosine similarity | CLS token versus the mean of its patch tokens | |
| Centered cosine reference | Means of two disjoint halves of the same patch-token set | |
| Top-128 projection | Patch token from the same image (included when fitting the directions) | |
| Top-128 projection | Patch token from an unrelated image | |
| Top-128 projection | CLS token from the same image | |
| Top-128 projection reference | Random 1,280-D vector ( expected) |
| Vector | Top-1 (%) | Variance kept | Relative accuracy |
|---|---|---|---|
| patch mean | 34.45 | 100% | — |
| CLS token | 18.95 | 100% | — |
| patch mean, CLS-token prediction removed | 10.80 | 24.2% | 31.3% |
| CLS token, patch-mean prediction removed | 2.00 | 35.2% | 10.6% |
| Feature family | Test Top-1 (%) | Train pairwise (%) |
|---|---|---|
| object category / semantics | 17.0 | 82 |
| color | 16.5 | 79 |
| texture | 14.5 | 78 |
| objectness / 2-D layout | 7.0 | 58 |
| HOG oriented edges | 6.0 | 67 |
| SAM 3 object silhouette | 6.5 | 55 |
| Brain encoder | Pretraining | Within Top-1 | LOSO Top-1 |
|---|---|---|---|
| EEG-Conformer | none | ||
| CBraMod | EEG corpus | ||
| CSBrain | EEG corpus | ||
| LaBraM | 2,500 h EEG | ||
| BIOT | EEG corpus | ||
| EEGProjection (our baseline) | none |
| Brain encoder | Within Top-1 | LOSO Top-1 |
|---|---|---|
| EEGProjection | ||
| TSConv (TSception) | ||
| ShallowConvNet | ||
| EEGNet | ||
| DeepConvNet |
| no projection | projection output dimension | |||||||
|---|---|---|---|---|---|---|---|---|
| 271 | 32 | 48 | 64 | 96 | 112 | 128 | 160 | |
| within Top-1 | ||||||||
| LOSO Top-1 | ||||||||
| Configuration | PixCorr | SSIM | AlexNet-2 | AlexNet-5 | Inception | CLIP | SwAV |
|---|---|---|---|---|---|---|---|
| RPA-only | 0.2245 | 0.4520 | 0.8599 | 0.9044 | 0.7702 | 0.8093 | 0.5212 |
| RPA + pooled-CLIP | 0.2251 | 0.4541 | 0.8645 | 0.9086 | 0.7832 | 0.8107 | 0.5256 |