cs.CVSep 19, 2026

RPA: Residual Patch-Token Adapter for Image Retrieval from EEG and MEG

Authors: Yuhui Jin, Yonghao Song, Bingchuan Liu

Organizations: Vortek Lab Inc. · Department of Computer Science and Technology, Tsinghua University · Biology and Biological Engineering, California Institute of Technology

Abstract

Most existing MEG and EEG (M/EEG) visual decoding methods align brain signals with a single global embedding extracted from a pretrained visual encoder, leaving open whether intermediate patch representations, which preserve richer and more granular rich visual information, can improve representation learning. To address this question, we introduce the Residual Patch Adapter (RPA), a lightweight, modular adapter that leverages all patch tokens from an intermediate layer of a ViT visual encoder for alignment. Through extensive ablation analyses, we first show that pooling or masking patch tokens degrades the learned representation, demonstrating that retaining the full set of patch tokens is important for EEG alignment, while the CLS token provides little unique information. We then use a series of six quantitative feature analyses to show that both higher-level semantics and lower-level visual features, including color and texture, are essential for this EEG-to-image alignment. Under current protocols, our system achieves Top-1 accuracies of 95.4% within-subject and 35.5% cross-subject on THINGS-EEG2, and 65.2% and 6.7%, respectively, on THINGS-MEG, achieving state-of-the-art (SOTA) performance across both datasets. Evaluations with alternative brain encoders, including pretrained EEG foundation models, demonstrate that the approach extends beyond the projection-based EEG encoder. Furthermore, we provide a plug-and-play interface that allows RPA to be replaced by convolution, attention, or ConvNeXt alternatives. Together, these findings provide significant insight into M/EEG-to-image representation learning by establishing design principles for leveraging the latent space of visual encoders, and open new directions for brain--image alignment and non-invasive brain--computer interface (BCI).

Figures & tables

Appendix figures & tables26 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MindAlign: Bridging EEG, Vision, and Language for Zero-Shot Visual Decoding

    May 23, 2026Zexuan Chen, Sichao Liu, Runhao Lu +4ElectroencephalographyNeural Recordings

  2. What Does the Brain See? Multiview Neural Representations to Demystify the Brain-Visual Alignment

    Jun 24, 2026Salini Yadav, Taveena Lotey, Pravendra Singh +1Electroencephalography DecodingVisual Perception