Spiking neural networks (SNNs) offer an energy-efficient computing paradigm through sparse event-driven computation, showing great potential for efficient multimodal learning. However, applying SNNs to high-level multimodal tasks, such as image-text retrieval (ITR), remains challenging, since sparse spike representations make it difficult to capture semantic structures required for cross-modal alignment. Existing spiking ITR methods rely on local alignment and additional soft-label supervision during training, while lacking awareness of structural and multi-granularity relationships. To address these issues, we propose a Multi-head Spiking Graph Attention Network (\textbf{MSGAT}) for structural modeling and equip it with dynamic attention heads to capture complementary relational patterns and enable spike-driven graph reasoning and aggregation. However, within a two-branch multi-granularity fusion framework, the fine-grained spike representations generated by MSGAT are sparse and discrete, whereas the global representations are continuous, making conventional feature-level fusion susceptible to interference across heterogeneous representations. Therefore, we introduce \textbf{Sim-Fuse}, a similarity-space fusion alignment strategy integrating coarse- and fine-grained matching relations while avoiding direct fusion of heterogeneous representations. Experiments on Flickr30K and MSCOCO show our method outperforms ANN methods under matched settings and existing SNN retrieval baselines. Moreover, with only two time steps, our SNN achieves comparable or superior performance to its ANN counterpart while reducing theoretical module-level energy by 55%. The code is provided in the Supplementary Materials.
Figures & tables
Figure 1: Motivation. Sparse, discrete SNN representations and continuous representations produced by ANN-based GPO (Generalized Pooling Operator) lie in heterogeneous spaces, making conventional feature-level fusion ineffective.
Figure 2: Framework of MSGAT for ITR. Left: Replaceable extractors produce floating-point node features. Top and bottom: GPO derives global features, Spike Encoder converts node features into spikes for Spiking Graph construction. MSGAT performs spike-driven graph reasoning and aggregation to obtain neighborhood-aware spiking fine-grained features. Middle: Coarse- and fine-grained similarities are fused into a detached multi-granularity target, which in turn aligns and optimizes the two branches.
Method
Venue
S
MSCOCO 1K (5-fold) Test Set
Flickr30K 1K Test Set
Image-to-Text
Text-to-Image
R@ Sum
Image-to-Text
Text-to-Image
R@ Sum
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
Faster R-CNN + BiGRU
SCAN ∗
ECCV18
72.7
94.8
98.4
58.8
88.4
94.8
507.9
67.4
90.3
95.8
48.6
77.7
85.2
465.0
VSRN ∗
ICCV19
76.2
94.8
98.2
62.8
89.7
95.1
516.8
71.3
90.6
96.0
54.7
81.8
88.2
482.6
VSE ∞
CVPR21
78.5
96.0
98.7
61.7
90.3
95.6
520.8
76.5
94.2
97.7
56.4
83.4
89.9
498.1
Table 1: Results on the MSCOCO and Flickr30K 1K test sets. Best and second-best results are marked in bold and underline , respectively. ∗ denotes ensemble results, and S indicates spike-driven methods. For fair comparison, we report only Faster R-CNN- and BERT/BiGRU-based baselines; results with pretrained VLMs are provided in the Supplementary Material.
Methods
MSCOCO 5K Test Set
Image-to-Text
Text-to-Image
R@ Sum
R@1
R@5
R@10
R@1
R@5
R@10
Faster R-CNN + BiGRU
SCAN ∗
50.4
82.2
90.0
38.6
69.3
80.4
410.9
VSRN ∗
53.0
81.1
89.4
40.5
70.6
81.1
415.7
VSE ∞
56.6
83.6
91.4
39.3
69.9
81.1
421.9
Table 2: Results on MSCOCO 5K Test Set. Best results are in bold, and the second-best scores are underlined . T=2 indicates the number of time steps in MSGAT, while ANN denotes the ANN counterpart of MSGAT.
Figure 4: Effectiveness of Similarity-Space Fuse strategy on Flickr30K 1K test set. Add-Fuse and Concat-Fuse denote features addition and concatenation, respectively.
Figure 5: Training Performance of MSGAT with 1 to 5 Time Steps. The upper set of curves corresponds to Sim-Fuse, whereas the lower set corresponds to without Sim-Fuse.
Fine-grained Branch
Image-to-Text
Text-to-Image
R@ Sum
R@1
R@5
R@10
R@1
R@5
R@10
w/o
82.5
95.0
98.0
61.5
85.0
90.7
512.7
SSA T=2
81.9
96.7
98.5
64.1
87.2
92.1
520.4
SGCN T=2
84.1
97.4
99.0
64.8
87.0
92.3
524.7
MSGAT T=2
86.1
98.4
99.0
65.9
87.8
92.6
529.8
GATv2-ANN
85.7
97.3
98.7
65.4
87.4
92.7
527.2
Table 3: Ablation of fine-grained branch on Flickr30K 1K test set. Best results are in bold.
Loss function
Image-to-Text
Text-to-Image
R@ Sum
Lf
Ln
Ls
LKL
R@1
R@5
R@10
R@1
R@5
R@10
✓
84.1
97.1
98.5
64.2
86.3
92.0
522.2
✓
✓
84.2
97.5
99.0
64.9
86.9
92.4
525.0
✓
✓
✓
85.6
97.8
99.0
64.8
87.0
92.2
526.4
✓
✓
✓
✓
86.1
98.4
99.0
65.9
87.8
92.6
529.8
Table 4: Influence of loss function on Flickr30K 1K test set. Best results are in bold.
Heads
Image-to-Text
Text-to-Image
R@ Sum
R@1
R@5
R@10
R@1
R@5
R@10
2
-
-
-
-
-
-
-
4
83.6
97.0
98.2
64.6
86.4
92.5
522.3
8
86.1
98.4
99.0
65.9
87.8
92.6
529.8
16
84.6
97.4
98.9
64.2
87.1
92.3
524.5
32
84.0
97.3
98.7
64.2
86.9
92.4
523.0
Table 5: Effect of the number of MSGAT dynamic attention heads on Flickr30K 1K test set. Best results are in bold.
Spiking Neural Networks (SNNs) enable event-driven computation with sparse activations, but building multimodal Transformers on SNNs is hindered by unstable training in deep spiking stacks and the mismatch between dense softmax attention and spike-based communication. We propose SMM Transformer, an SNN-based multimodal Transformer framework that combines (i)PLMP, a Parallel LIF with Multistage Learnable Parameters neuron and a tailored P-STBP algorithm for stable deep SNN training, (ii) SMSA, an attention-inspired spike-driven token-mixing module that replaces dense pairwise softmax attention with channel-wise spike co-activation and self-compensation, and (iii)SMoE, a spiking mixture-of-experts module for modality-aware fusion. Across visual and multimodal benchmarks, SMM Transformer achieves competitive accuracy compared to ANN baselines. Under a standard MAC/AC arithmetic model, SMSA reduces the estimated operator-level compute energy of the attention module by up to 97%, while whole-model profiling shows more moderate but consistent efficiency gains.
Xiubo Liang, Jinxing Han, Yuke Li +3
School of Software Technology, Zhejiang University, Ningbo, China · NetEase Yidun AI Lab, Hangzhou, China
We propose a new spiking neural network (SNN) design to process static images and event streams using time-to-first-spike (TTFS) latencies. Our key research question is which architectural choices best accommodate local and online learning in multi-layer convolutional SNNs. This question is addressed via an original framework combining residual-like connections with multi-depth feature aggregation and consensus. The full SNN pipeline features an early-vision front end, to convert raw visual data into sparse spike latencies, a four-layer convolutional backbone trained layerwise with unsupervised spike-timing-dependent plasticity (STDP), a deterministic Multi-Depth Temporal Fusion (MDTF) and a final classifier trained with reward-modulated spike-timing-dependent plasticity (R-STDP). Rather than replacing early features in deeper layers, the proposed MDTF preserves early temporal evidence, adding sparse residual events from intermediate layers, and incorporating deeper features only when they agree in time with earlier representations. The resulting architecture is experimentally validated across MNIST, Fashion-MNIST, CIFAR-10, and N-MNIST, delivering strong classification performance under a fully local learning regime. Selective multi-depth fusion significantly outperforms traditional STDP/R-STDP baselines on higher-variability visual tasks (achieving +18.2 pp on Fashion-MNIST and +29.2 pp on CIFAR-10). Furthermore, activity-budget analyses show that the network retains high accuracy even when removing a large fraction of late or weak spike events, confirming its high data efficiency and reduced event-processing requirements. The codebase is publicly available at github.com/aidinattar/multi-depth-temporal-fusion-snn.
Aidin Attar, Eleonora Cicciarella, Michele Rossi
Department of Information Engineering, University of Padua Padua, Italy
Infrared and visible image fusion (IVIF) integrates the complementary information of two modalities into a single image with richer scene content. While existing methods are largely built on artificial neural networks (ANNs), which densely compute over all activations, spiking neural networks (SNNs) communicate through sparse binary spikes and compute only where and when a spike occurs, offering a route to more energy-efficient fusion. However, directly applying SNNs to IVIF creates a fundamental tension: cross-modal fusion relies on fine-grained responses from both modalities, whereas binary spikes can discard complementary cues that remain below the firing threshold. The membrane potential retains these subthreshold responses before firing, letting both modalities jointly shape the output when integrated at this stage. Building on this, we propose CIS-Fuse, a spiking network that performs cross-modal fusion directly at the membrane-potential level. At its core is the current injection spiking (CIS) operator, which injects one modality as a gated auxiliary current into the driving neuron of the other, so the two integrate before spike firing, with a per-channel learnable injection strength that adaptively regulates the modulation magnitude. Building on CIS, we construct a bidirectional cross-modal fusion (BCMF) module and deploy it on a dual-branch architecture with asymmetric stacking depths, where the two branches develop a clear functional specialization. Extensive experiments on four IVIF benchmarks and on downstream detection and segmentation show that CIS-Fuse achieves fusion quality on par with state-of-the-art ANN-based methods while inheriting the energy efficiency of spike-based computation, with roughly an order of magnitude lower inference energy than the similarly-sized ANN-based DCEvo. Code will be released upon publication.