RPA: Residual Patch-Token Adapter for Image Retrieval from EEG and MEG
Authors: Yuhui Jin, Yonghao Song, Bingchuan Liu
Organizations: Vortek Lab Inc. · Department of Computer Science and Technology, Tsinghua University · Biology and Biological Engineering, California Institute of Technology
Most existing MEG and EEG (M/EEG) visual decoding methods align brain signals with a single global embedding extracted from a pretrained visual encoder, leaving open whether intermediate patch representations, which preserve richer and more granular rich visual information, can improve representation learning. To address this question, we introduce the Residual Patch Adapter (RPA), a lightweight, modular adapter that leverages all patch tokens from an intermediate layer of a ViT visual encoder for alignment. Through extensive ablation analyses, we first show that pooling or masking patch tokens degrades the learned representation, demonstrating that retaining the full set of patch tokens is important for EEG alignment, while the CLS token provides little unique information. We then use a series of six quantitative feature analyses to show that both higher-level semantics and lower-level visual features, including color and texture, are essential for this EEG-to-image alignment. Under current protocols, our system achieves Top-1 accuracies of 95.4% within-subject and 35.5% cross-subject on THINGS-EEG2, and 65.2% and 6.7%, respectively, on THINGS-MEG, achieving state-of-the-art (SOTA) performance across both datasets. Evaluations with alternative brain encoders, including pretrained EEG foundation models, demonstrate that the approach extends beyond the projection-based EEG encoder. Furthermore, we provide a plug-and-play interface that allows RPA to be replaced by convolution, attention, or ConvNeXt alternatives. Together, these findings provide significant insight into M/EEG-to-image representation learning by establishing design principles for leveraging the latent space of visual encoders, and open new directions for brain--image alignment and non-invasive brain--computer interface (BCI).
Figures & tables
Figure 1: Overview of the EEG-to-image retrieval framework. (A) EEG is recorded while a participant views an image. (B) A frozen visual encoder and a trainable visual adapter produce the visual embedding, which is aligned with the brain embedding for retrieval. (C) The schematic contrasts a final-layer CLS route with an intermediate-patch route; the controlled comparison in Table 3 uses CLS and patch tokens from the same intermediate layer. (D) The residual adapter branch transforms the patch-token grid with convolution–normalization–GELU layers. Pooling and projection, shown separately in panel B, complete the visual adapter.
EEG
MEG
Within
Cross
Within
Cross
Method
T1
T5
T1
T5
T1
T5
T1
T5
NICE ( Song et al., 2024 )
13.8
39.5
6.2
21.4
12.8
36.0
–
–
ATM ( Li et al., 2024 )
28.6
58.5
11.8
33.7
–
–
–
–
UBP ( Wu et al., 2025 )
50.9
79.7
12.4
33.4
26.7
55.2
2.2
10.4
BrainHIVE ( Zheng et al., 2026 )
75.7
94.6
20.0
44.1
33.7
60.5
5.4
15.2
Table 1: Average Top-1/Top-5 accuracy (%) for 200-way zero-shot retrieval on THINGS-EEG2 and THINGS-MEG. Values are participant averages; “–” indicates not reported.
Low-level
High-level
Method
PixCorr ↑
SSIM ↑
AlexNet-2 ↑
AlexNet-5 ↑
Inception ↑
CLIP ↑
SwAV ↓
CognitionCapturer (All) ( Zhang et al., 2025 )
0.150
0.347
0.754
0.623
0.669
0.715
0.590
AVDE ( Dai et al., 2026 )
0.147
0.366
0.766
0.835
0.724
0.747
0.587
BrainHIVE ( Zheng et al., 2026 )
0.195
0.336
0.843
0.905
0.756
0.808
0.554
Ours
0.225
0.454
0.865
0.909
0.783
0.811
0.526
Table 2: Average image-reconstruction performance on THINGS-EEG2.
Figure 2: Visual-encoder depth profiles for EEG and MEG under within-subject and LOSO evaluation. The x-axis shows six uniformly spaced relative depths. Curves show mean 200-way Top-1 scores, averaged across three independent runs and standardized within each encoder.
Visual input
Visual adapter
Top-1 (%)
CLS token
LayerNorm + linear
69.50±0.45
patch tokens
mean, LayerNorm + linear
77.03±0.7
patch tokens
pointwise RPA
92.98±0.2
patch tokens
local RPA
95.4±0.1
Table 3: Controlled within-subject THINGS-EEG2 comparison at a fixed visual-encoder output.
Figure 3: RPA diagnostics. (A–B) Top-1 accuracy for native, shuffled, and spatially averaged patch-token grids. (C–D) Top-1 accuracy under increasing zero, CLS-token, or Gaussian replacement. Solid gray lines mark unmasked performance.
Figure 4: Qualitative EEG-to-image reconstructions. Each row shows one animal concept: the viewed image followed by reconstructions from Subjects 1–10.
Figure 5: Alignment, timing, averaging, and attribution diagnostics. (A) UMAP projection of EEG embeddings; colors indicate categories, and insets show example stimuli. (B) Given layers of visual encoder and 100-ms EEG windows advanced in 60-ms steps, we using a ridge probe to obtain a decoding accuracy. (C) Retrieval with ridge-predicted means of frozen patch-token subsets; the dashed line marks the CLS-token result. (D) Top-1 accuracy versus the number of EEG test trials averaged per stimulus; dark and light blue denote within-subject and LOSO evaluation. (E) Grad-CAM maps; columns show the stimulus, image-only RPA, and RPA.
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Within subject
LOSO
Training responses
16,540 from one subject
148,860 from nine subjects
Test subject
same subject, unseen concepts
held-out subject, unseen concepts
Default window
0–700 ms
0–1,000 ms
Epochs
25
15
Peak learning rate
5×10−4
3×10−4
Projection dropout
0.5
0.7
Appendix
Table 4: Within-subject and LOSO EEG training configurations.
Method
S1
S2
S3
S4
S5
S6
S7
S8
S9
S10
Mean
Top-1 accuracy (%)
NICE ( Song et al., 2024 )
12.3
10.4
13.1
16.4
8.0
14.1
15.2
20.0
13.3
14.9
13.8
UBP ( Wu et al., 2025 )
41.2
51.2
51.2
51.1
42.2
57.5
49.0
58.6
45.1
61.5
50.9
ViEEG ( Liu et al., 2026a )
34.1
38.4
40.6
50.1
28.9
44.3
38.6
54.0
37.3
42.8
40.9
NeuroBridge ( Zhang et al., 2026 )
50.0
63.2
61.6
61.4
54.8
69.7
62.7
71.2
64.0
73.6
63.2
BrainHIVE ( Zheng et al., 2026 )
64.3
76.3
74.0
67.0
68.0
81.5
76.8
84.8
76.8
87.3
75.7
Appendix
Table 5: Subject-level within-subject retrieval on THINGS-EEG2.
Method
S1
S2
S3
S4
S5
S6
S7
S8
S9
S10
Mean
Top-1 accuracy (%)
NICE ( Song et al., 2024 )
7.6
5.9
6.0
6.3
4.4
5.6
5.6
6.3
5.7
8.4
6.2
UBP ( Wu et al., 2025 )
11.5
15.5
9.8
13.0
8.8
11.7
10.2
12.2
15.5
16.0
12.4
ViEEG ( Liu et al., 2026a )
22.7
24.7
19.0
25.5
19.8
20.7
20.9
20.8
23.8
31.2
22.9
NeuroBridge ( Zhang et al., 2026 )
23.2
21.2
13.2
17.0
14.5
25.0
15.3
20.1
13.7
27.2
19.0
BrainHIVE ( Zheng et al., 2026 )
–
–
–
–
–
–
–
–
–
–
20.0
Appendix
Table 6: Subject-level LOSO retrieval on THINGS-EEG2.
Method
S1
S2
S3
S4
Mean
Top-1 accuracy (%)
NICE ( Song et al., 2024 )
9.6
18.5
14.2
9.0
12.8
UBP ( Wu et al., 2025 )
15.0
46.0
27.3
18.5
26.7
ViEEG ( Liu et al., 2026a )
16.6
37.4
30.6
17.2
25.5
NeuroBridge ( Zhang et al., 2026 )
16.5
53.7
40.4
18.1
32.2
BrainHIVE ( Zheng et al., 2026 )
14.0
63.8
41.0
17.0
33.9
Appendix
Table 7: Subject-level within-subject retrieval on THINGS-MEG.
Method
S1
S2
S3
S4
Mean
Top-1 accuracy (%)
UBP ( Wu et al., 2025 )
2.0
1.5
2.7
2.5
2.2
NeuroBridge ( Zhang et al., 2026 )
4.3
3.6
3.0
2.5
3.4
BrainHIVE ( Zheng et al., 2026 )
–
–
–
–
5.4
Blur ( Liu et al., 2026b )
2.9
7.7
5.8
4.7
5.3
Shallow Align. ( Du et al., 2026 )
1.3
6.6
5.4
1.5
3.7
Appendix
Table 8: Subject-level LOSO retrieval on THINGS-MEG.
Family
Adapter
Within
LOSO
RPA
local (1,3,1)
95.4
35.5
RPA
pointwise (1,1,1)
93.0
34.9
RPA ablation
no final GELU
90.9
32.4
RPA ablation
pointwise + query pool
90.7
34.0
scalar attention
residual-MLP values
89.7
32.5
scalar attention
MLP values
89.2
31.5
Appendix
Table 9: Comparison of visual adapters (200-way Top-1, %).
Adapter
Parameters
Adapter
Parameters
CLS-token / patch-mean controls
1.314M
patch mean + linear (no LN)
1.312M
pointwise RPA (1,1,1)
2.038M
local RPA (1,3,1)
2.562M
local RPA, no final GELU
2.562M
depthwise-separable conv
2.969M
query pool + MLP values
3.610M
pointwise RPA + query pool
3.679M
query pool + linear values
4.592M
8-head MHA query pool
7.872M
full-width ConvNeXt
14.493M
transformer + MHA pool
20.993M
Appendix
Table 10: Trainable visual-adapter parameters, each maps visual embedding of 1,024 dimensions. Counts are for the fixed image encoder used in Table 9 .
Adapter
Adapter output hv before L2 normalization
local RPA (1,3,1)
Q(GAP[g+B131(g)])
pointwise RPA (1,1,1)
Q(GAP[g+B111(g)])
local RPA, no final GELU
local RPA with the last GELU in B131 removed; the residual +g is retained
Q(AttnPoolVmlp(t)) , Vmlp(x)=W2GELU(W1x) with the same widths
Appendix
Table 11: Definitions of the visual adapters evaluated in Table 9 .
Adapter
Structure
Params
Within subject
LOSO
width 64
r=64
1.52M
93.4
34.9
no local mixing
k=(1,1,1)
2.04M
93.0
34.8
reference
r=256,k=(1,3,1)
2.56M
95.4
35.5
all-local
k=(3,3,3)
7.81M
91.9
33.8
width 1280
r=1280
19.35M
94.5
35.3
Appendix
Table 12: RPA bottleneck-width and kernel-placement study.
Encoder
rank 1
rank 2
rank 3
rank 4
rank 5
rank 6 (final)
EEG within-subject
CLIP-H
93.0±0.2
94.0±0.1
92.2±0.2
69.3±0.9
61.4±0.1
52.7±0.8
I-JEPA
85.7±0.4
89.9±0.5
91.9±0.2
90.6±0.6
89.9±0.1
73.4±0.6
DINOv2-L
89.1±0.4
92.9±0.2
94.4±0.5
90.8±0.6
77.8±0.5
55.8±0.6
DINOv2-S
88.7±0.3
91.0±0.1
89.5±0.4
84.3±0.4
77.0±0.3
62.2±1.3
ViT-B/32
91.7±0.4
93.8±0.2
91.3±0.2
77.1±0.7
56.2±1.2
51.7±0.8
Appendix
Table 13: Full visual-encoder layer sweep (EEGProjection encoder, uniform six-point grid; Top-1 accuracy in %). Rank 1 is the shallowest sampled state and rank 6 is the final block.
Encoder
Pretraining
Final layer
Test-best sampled layer (rank)
CLIP ViT-H/14
language contrastive
19.3±0.9
32.9±0.4 (r3)
I-JEPA ViT-H/14
masked SSL
25.7±0.6
31.6±0.5 (r3)
DINOv2-L/14
SSL distillation
19.5±0.6
31.9±0.4 (r3)
DINOv2-S/14
SSL distillation
20.7±0.3
29.3±0.8 (r2)
CLIP ViT-B/32
language contrastive
18.4±0.2
31.6±0.3 (r2)
CLIP RN50
language contrastive
25.0±0.3
28.2±1.2 (r5)
Appendix
Table 14: LOSO EEG comparison between each encoder’s final sampled state and the maximum over its six test-set scores.
Analysis
Vector or comparison
Result
Centered cosine similarity
CLS token versus the mean of its patch tokens
−0.05
Centered cosine reference
Means of two disjoint halves of the same patch-token set
0.89
Top-128 projection
Patch token from the same image (included when fitting the directions)
0.92
Top-128 projection
Patch token from an unrelated image
0.36
Top-128 projection
CLS token from the same image
0.17
Top-128 projection reference
Random 1,280-D vector ( 128/1280 expected)
0.10
Appendix
Table 15: Three comparisons between the frozen CLS token and patch tokens on the 200 test images. Before computing cosine similarity, we subtract the across-image mean separately from the CLS tokens and patch means. The projection rows report the fraction of a vector’s squared norm captured by the top 128 directions fitted to the same image’s patch tokens. These are deterministic visual-feature summaries; no EEG or RPA is involved.
Vector
Top-1 (%)
Variance kept
Relative accuracy
patch mean
34.45
100%
—
CLS token
18.95
100%
—
patch mean, CLS-token prediction removed
10.80
24.2%
31.3%
CLS token, patch-mean prediction removed
2.00
35.2%
10.6%
Appendix
Table 16: Reciprocal residualization followed by the same linear EEG probe. The residualization map is fit on training images and applied unchanged to test images. Relative accuracy is Top-1 accuracy as a percentage of that for the corresponding unmodified target, without chance correction.
Feature family
Test Top-1 (%)
Train pairwise (%)
object category / semantics
17.0
82
color
16.5
79
texture
14.5
78
objectness / 2-D layout
7.0
58
HOG oriented edges
6.0
67
SAM 3 object silhouette
6.5
55
Appendix
Table 17: Exploratory linear decoding of six named image-feature targets from EEG (10 participants). Test Top-1 is five-fold, 40-way cosine retrieval on the 200-image test set (chance 2.5% ). Train pairwise is matched-versus-mismatched accuracy on up to 8,000 training images (chance 50% ).
Brain encoder
Pretraining
Within Top-1
LOSO Top-1
EEG-Conformer
none
94.90±0.28
34.45±0.26
CBraMod
EEG corpus
94.55±0.40
35.63±0.06
CSBrain
EEG corpus
93.83±0.36
28.73±0.61
LaBraM
2,500 h EEG
79.33±1.10
32.22±0.61
BIOT
EEG corpus
32.57±0.39
11.47±1.29
EEGProjection (our baseline)
none
95.4±0.12
35.5±0.34
Appendix
Table 18: Within-subject and LOSO 200-way Top-1 accuracy (%) EEG foundation models.
Brain encoder
Within Top-1
LOSO Top-1
EEGProjection
95.4±0.3
35.5±0.3
TSConv (TSception)
87.9±0.6
30.4±0.5
ShallowConvNet
79.0±0.8
23.9±0.6
EEGNet
77.2±0.6
20.7±0.8
DeepConvNet
64.7±0.1
15.3±0.4
Appendix
Table 19: Within-subject and LOSO Top-1 accuracy (%) for various EEG encoders
Figure 6: Top-1 accuracy for matched-duration EEG windows retained from the start, [0,d] , or end, [1000−d,1000] .
Figure 7: Error-focused Top-10 retrievals for Subject 1 on THINGS-EEG2: seven rows are selected Top-1 errors and one is a Top-1 success, while the complete test-set Top-1 accuracy is 96% ; red outlines identify the viewed image and its position among the ranked candidates.
Figure 8: EEG–image cosine-similarity matrix averaged over 10 within-subject models. Rows are EEG embeddings, columns are visual embeddings, and the diagonal contains the correct pairs.
no projection
projection output dimension K
271
32
48
64
96
112
128
160
within Top-1
51.3±0.5
58.7±1.1
59.2±0.7
65.2±1.0
64.0±3.0
63.4±2.1
64.2±1.8
63.2±1.5
LOSO Top-1
5.1±0.7
5.8±0.2
6.2±0.5
6.7±0.4
5.9±0.4
6.4±0.2
6.1±0.2
6.0±0.4
Appendix
Table 20: Effect of the MEG sensor-projection output dimension. “No projection” uses all 271 sensors directly.
Configuration
PixCorr ↑
SSIM ↑
AlexNet-2 ↑
AlexNet-5 ↑
Inception ↑
CLIP ↑
SwAV ↓
RPA-only
0.2245
0.4520
0.8599
0.9044
0.7702
0.8093
0.5212
RPA + pooled-CLIP
0.2251
0.4541
0.8645
0.9086
0.7832
0.8107
0.5256
Appendix
Table 21: Reconstruction results.
Figure 9: Viewed images (top) and participant-level reconstructions from 80-repetition-averaged EEG responses for eight animal stimuli.
Figure 10: Viewed images (top) and participant-level reconstructions from 80-repetition-averaged EEG responses for eight vegetable and herb stimuli.
Figure 11: Viewed images (top) and participant-level reconstructions from 80-repetition-averaged EEG responses for eight baked-good and dessert stimuli.
Figure 12: Viewed images (top) and participant-level reconstructions from 80-repetition-averaged EEG responses for eight clothing stimuli.
Figure 13: Viewed images (top) and participant-level reconstructions from 80-repetition-averaged EEG responses for eight vehicle stimuli.
Visual decoding from brain signals is a key challenge at the intersection of computer vision and neuroscience, requiring methods that bridge neural representations and computational models of vision. We introduce a tri-modal contrastive framework for EEG-based visual decoding that aligns EEG, visual, and textual representations within a unified latent space. Our approach follows a two-stage design. First, we pre-train an EEG encoder via masked reconstruction on unlabeled trials, learning spatio-temporal regularities that transfer robustly to downstream tasks. Second, we jointly align EEG, image, and LLM-generated textual descriptions through contrastive learning, where text supervision acts as a semantic regularizer that injects linguistic structure into the shared space without overwhelming the primary EEG-image signal. The encoder integrates subject-specific adaptation, graph-attention over channels, and temporal-spatial convolutional embeddings. On the Things-EEG2 200-way zero-shot benchmark, our framework achieves 54.1% Top-1 and 83.4% Top-5 accuracy, substantially exceeding the strongest prior baseline (32.4% / 64.0%), with paired Wilcoxon tests confirming significance (p < 0.01) over all in-subject baselines. We validate generalization on Things-MEG. Analysis reveals that compact embedding geometries (CN-CLIP) outperform much larger backbones, and that decoding aligns with established neurophysiology of visual processing. This work is a critical step towards robust, semantically-grounded visual decoding from non-invasive temporal neural signals. The source code is publicly available in https://github.com/anon-eeg/eeg_image_decoding.
Zexuan Chen, Sichao Liu, Runhao Lu +4
KTH, SWeden · University of Cambridge, UK · EPFL, Switzerland +2
Zero-shot visual decoding from electroencephalography (EEG) aims to infer visual semantics from non-invasive neural recordings, but remains challenging due to the low signal-to-noise ratio, non-stationarity, and limited spatial resolution of EEG. Existing EEG-vision alignment methods often rely on holistic EEG embeddings, which can obscure the complementary temporal, spectral, and spatial structure underlying visual perception. We introduce a unified multiview EEG representation learning framework for aligning brain responses with visual semantic embeddings. Our method builds an EEG encoder that jointly models three complementary views: input-conditioned state-space temporal dynamics, learnable wavelet-based spectral decomposition for sample-adaptive frequency modeling, and attention-modulated graph learning for structured electrode interactions. The resulting multiview EEG embeddings are fused and aligned with pretrained visual representations in a shared semantic space using contrastive learning with EEG-specific regularization, enabling 200-way zero-shot visual classification. Experiments on THINGS-EEG benchmark show that our method achieves state-of-the-art performance, with 54.8% Top-1 and 85.6% Top-5 accuracy in the within-subject setting and 15.3% Top-1 and 45.4% Top-5 accuracy in the cross-subject setting. We further present the first systematic cross-session EEG-image decoding evaluation, achieving 40.8% Top-1 and 78.0% Top-5 accuracy. These results suggest that explicitly modeling multiview neural structure improves both semantic alignment and generalization in EEG-based visual decoding.
Salini Yadav, Taveena Lotey, Pravendra Singh +1
Department of CSE IIT Roorkee, India · Department of CSE IIT (ISM) Dhanbad, India
We present a brain-to-image system that decodes visual stimuli from EEG signals recorded during natural image viewing. Our system addresses two tasks: (1) EEG-to-image retrieval, which ranks the correct stimulus image among 200 candidates given an EEG segment, and (2) EEG-to-image reconstruction, which generates an image consistent with the perceived stimulus. For retrieval, we implement a multi-level blurring approach improved with biologically inspired EVNet features and trained with the InfoNCE loss. Evaluated over 10 random seeds for a single subject, the retrieval model achieves a mean final-epoch Top-1 accuracy of 86.30% and Top-5 accuracy of 98.55%. For reconstruction, we implement CognitionCapturerPro, which aligns EEG representations to multi-modal CLIP embeddings, including image, text, depth, and edge embeddings, and synthesizes images with SDXL-Turbo conditioned via IP-Adapter. Averaged over 10 seeds, the reconstruction model achieves a CLIP score of 0.903 using ViT-H-14, a CLIP score of 0.870 using ViT-L/14, and an SSIM of 0.409. These results demonstrate the feasibility of decoding rich visual representations from EEG signals using modern multi-modal alignment and generative modeling techniques.