Cross-subject EEG-to-image retrieval requires a neural represen- tation trained on source subjects to remain aligned with a visual embedding space for an unseen subject. Whereas existing methods primarily focus on the EEG side, we address this problem from the perspective of the visual target. Our approach preserves the spatial information of the Perception Encoder, converts its patch grid into a compact set of learned visual views, and aggregates them for each image with a block-structured, content-dependent router. The target is learned jointly with the EEG encoder through contrastive learning with MMD regularization across source subjects. For deployment, we propose a training-free representation refinement that aligns frozen embeddings without updating either encoder. Under leave- one-subject-out evaluation on THINGS-EEG2, the structured target achieves 35.3%/65.6% Top-1/Top-5 accuracy, the best among com- pared methods. Refinement raises this to 48.1%/77.1%, an 18.5% Top-1 gain over the strongest compared method, improving all ten held-out subjects.
Figures & tables
Figure 1: Overview of the proposed cross-subject EEG-to-image retrieval framework. The trainable EEG encoder maps EEG signals into the shared embedding space, while a frozen pre-trained vision encoder produces spatial visual tokens that are converted into multiple visual views using learned multi-view pooling and fused by the block attention–residual router. EEG and image embeddings are aligned through a two-stage objective, combining contrastive learning with MMD-based distribution alignment in Stage I and contrastive learning alone in Stage II.
Method
Sub-01
Sub-02
Sub-03
Sub-04
Sub-05
Sub-06
Sub-07
Sub-08
Sub-09
Sub-10
Avg.
NICE [ 17 ]
7.6/22.8
5.9/20.5
6.0/22.3
6.3/20.7
4.4/18.3
5.6/22.2
5.6/19.7
6.3/22.0
5.7/17.6
8.4/28.3
6.2/21.4
ATM [ 11 ]
10.5/26.8
7.1/24.8
11.9/33.8
14.7/39.4
7.0/23.9
11.1/35.8
16.1/43.5
15.0/40.3
4.9/22.7
20.5/46.5
11.9/33.8
UBP [ 21 ]
11.5/29.7
15.5/40.0
9.8/27.0
13.0/32.3
8.8/33.8
11.7/31.0
10.2/23.8
12.2/32.2
15.5/40.5
16.0/43.5
12.4/33.4
NeuroBridge [ 22 ]
23.2/52.4
21.2/49.3
13.2/36.5
17.0/45.3
14.5/37.7
25.0/55.0
15.3/45.1
20.1/44.9
13.7/36.5
27.2/56.3
19.0/45.9
Shallow Alignment [ 2 ]
24.6/54.7
31.3/61.5
11.4/31.1
19.9/48.8
19.0/45.5
24.1/49.8
18.6/51.6
17.6/46.7
23.3/54.9
34.6/63.2
22.4/50.8
SAMGA ∗ [ 9 ]
36.0/58.5
37.5/68.5
21.0/43.0
29.0/56.0
21.0/51.0
32.0/62.0
25.5/55.5
26.0/50.0
24.0/51.0
43.5/72.5
29.6/56.8
Table 1: Cross-subject zero-shot EEG-to-image retrieval on THINGS-EEG2 under leave-one-subject-out evaluation. Results are reported as Top-1 / Top-5 accuracy (%).
Figure 2: Subject-wise Top-1 retrieval accuracy on THINGS-EEG2 under leave-one-subject-out evaluation. Results are reported for each of the ten held-out subjects across all compared methods.
Figure 3: Average cross-subject Top-1 retrieval accuracy on THINGS-EEG2 under leave-one-subject-out evaluation. Results are averaged across the ten held-out subjects for all compared methods.
Figure 4: Subject-wise effect of representation refinement on cross-subject EEG-to-image retrieval. (a) Top-1 and Top-5 retrieval accuracy before and after refinement for each held-out subject. (b) Subject-wise Top-1 and Top-5 accuracy gains, reported in percentage points (pp).
Zero-shot EEG-to-image retrieval aims to decode perceived visual content from electroencephalography (EEG) by aligning neural responses with pretrained visual representations, providing a promising route toward scalable visual neural decoding and practical brain-computer interfaces. However, robust EEG-to-image retrieval remains challenging, because prior methods usually rely on either a single fixed visual target or a subject-invariant target construction scheme. Such designs overlook two important properties of visually evoked EEG signals: they preserve information across multiple representational scales, and the visual granularity best matched to EEG may vary across subjects. To address these issues, subject-aware multi-granularity alignment (SAMGA) framework is proposed for zero-shot EEG-to-image retrieval. SAMGA first constructs a subject-aware visual supervision target by adaptively aggregating multiple intermediate representations from a pretrained vision encoder, allowing the model to absorb subject-dependent granularity deviations during training while preserving subject-agnostic inference. Building on this adaptive target construction, a coarse-to-fine cross-modal alignment strategy is further designed with a shared encoder wherein the coarse stage stabilizes the shared semantic geometry and reduces subject-induced distribution shift, and the fine stage further improves instance-level retrieval discrimination. Extensive experiments on the THINGS-EEG benchmark demonstrate that the proposed method achieves 91.3% Top-1 and 98.8% Top-5 accuracy in the intra-subject setting, and 34.4% Top-1 and 64.8% Top-5 accuracy in the inter-subject setting, outperforming recent state-of-the-art methods.
Lin Jiang, Qingshan She, Jiale Xu +3
School of Automation, Hangzhou Dianzi University, and also with Zhejiang Provincial Key Laboratory of Brain Computer Collaborative Intelligence Technology and Applications, Hangzhou, Zhejiang 310018, China
Zero-shot brain-to-image retrieval requires robust alignment between noisy EEG responses and visual representations. Existing EEG-vision alignment methods often operate in sensor space and apply fixed visual supervision to all responses, ignoring both spatial mixing in scalp EEG and response-wise variability in alignment reliability. We propose an adaptive cortically constrained EEG-vision alignment method for zero-shot brain-to-image retrieval. The method reconstructs EEG responses into predefined ROI-level source-pattern representations and encodes them with a Neuro-ROI Attention Encoder. To handle response-wise variability, we introduce an evidence-based adaptive visual supervision strategy that weights detail-controlled visual targets using model-based alignment evidence. On THINGS-EEG, the proposed method achieves strong 200-way zero-shot retrieval performance, with ROI-level attribution providing post hoc interpretability of the learned source-pattern representations. These results show that cortically constrained representation learning and adaptive supervision can jointly support EEG-vision alignment for zero-shot brain-to-image retrieval.
Ye Wang, Haokun Ren, Wei Wu +4
School of Artificial Intelligence, Chongqing University of Posts and Telecommunications, Chongqing, China · School of Medicine, Shanghai Jiaotong University, Shanghai, China · National Center for Applied Mathematics in Chongqing, Chongqing Normal University, Chongqing, China +1
Zero-shot visual decoding from electroencephalography (EEG) aims to infer visual semantics from non-invasive neural recordings, but remains challenging due to the low signal-to-noise ratio, non-stationarity, and limited spatial resolution of EEG. Existing EEG-vision alignment methods often rely on holistic EEG embeddings, which can obscure the complementary temporal, spectral, and spatial structure underlying visual perception. We introduce a unified multiview EEG representation learning framework for aligning brain responses with visual semantic embeddings. Our method builds an EEG encoder that jointly models three complementary views: input-conditioned state-space temporal dynamics, learnable wavelet-based spectral decomposition for sample-adaptive frequency modeling, and attention-modulated graph learning for structured electrode interactions. The resulting multiview EEG embeddings are fused and aligned with pretrained visual representations in a shared semantic space using contrastive learning with EEG-specific regularization, enabling 200-way zero-shot visual classification. Experiments on THINGS-EEG benchmark show that our method achieves state-of-the-art performance, with 54.8% Top-1 and 85.6% Top-5 accuracy in the within-subject setting and 15.3% Top-1 and 45.4% Top-5 accuracy in the cross-subject setting. We further present the first systematic cross-session EEG-image decoding evaluation, achieving 40.8% Top-1 and 78.0% Top-5 accuracy. These results suggest that explicitly modeling multiview neural structure improves both semantic alignment and generalization in EEG-based visual decoding.
Salini Yadav, Taveena Lotey, Pravendra Singh +1
Department of CSE IIT Roorkee, India · Department of CSE IIT (ISM) Dhanbad, India