cs.CVOct 8, 2026

MSGAT: Multi-Head Spiking Graph Attention with Similarity-Space Fusion for Image-Text Retrieval

Authors: Xintao Zong, Wenxuan Liu, Jianhao Ding, Zhaofei Yu, Tiejun Huang

Organizations: Peking University

Abstract

Spiking neural networks (SNNs) offer an energy-efficient computing paradigm through sparse event-driven computation, showing great potential for efficient multimodal learning. However, applying SNNs to high-level multimodal tasks, such as image-text retrieval (ITR), remains challenging, since sparse spike representations make it difficult to capture semantic structures required for cross-modal alignment. Existing spiking ITR methods rely on local alignment and additional soft-label supervision during training, while lacking awareness of structural and multi-granularity relationships. To address these issues, we propose a Multi-head Spiking Graph Attention Network (\textbf{MSGAT}) for structural modeling and equip it with dynamic attention heads to capture complementary relational patterns and enable spike-driven graph reasoning and aggregation. However, within a two-branch multi-granularity fusion framework, the fine-grained spike representations generated by MSGAT are sparse and discrete, whereas the global representations are continuous, making conventional feature-level fusion susceptible to interference across heterogeneous representations. Therefore, we introduce \textbf{Sim-Fuse}, a similarity-space fusion alignment strategy integrating coarse- and fine-grained matching relations while avoiding direct fusion of heterogeneous representations. Experiments on Flickr30K and MSCOCO show our method outperforms ANN methods under matched settings and existing SNN retrieval baselines. Moreover, with only two time steps, our SNN achieves comparable or superior performance to its ANN counterpart while reducing theoretical module-level energy by 55%. The code is provided in the Supplementary Materials.

Figures & tables

Explore similar work

CardsList
  1. SMM Transformer: Leveraging Spiking Neural Networks for Multimodal Tasks

    Aug 3, 2026Xiubo Liang, Jinxing Han, Yuke Li +3Spiking TransformersSpiking Neural Networks

  2. Multi-Depth Temporal Fusion for Feedforward, Locally Trained Spiking Neural Networks

    Sep 29, 2026Aidin Attar, Eleonora Cicciarella, Michele RossiSpiking Neural NetworksFeature Fusion