Contrastive Attention Mitigates Spectral Bias in Spiking Transformers
Organizations: School of Computer Science and Engineering, University of Electronic Science and Technology of China
Abstract
Spiking Transformers merge the energy-efficiency of spiking neural networks (SNNs) with the representational power of self-attention, creating a promising architecture for high-performance, energy-efficient computation. However, a performance gap persists versus its counterparts in artificial neural networks (ANNs). Unlike prior works attributing this to binary activations, we reveal that both spiking neurons and spiking self-attention (SSA) act as low-pass filters through multiscale spectral analysis. This characteristic leads to the dissipation of high-frequency components. To address this issue, we propose the Spiking Contrastive Attention (SCA) paradigm, which draw inspiration from the edge-detection and differential sensing properties of biological visual system. By extracting contrast prototypes via global contrastive aggregation and applying local differential refinement, SCA effectively enhances high-frequency information. Extensive experiments show that SCA is a general module that consistently boosts Spiking Transformers across image classification, semantic segmentation, and event-based tracking. Furthermore, it achieves lower complexity, offering superior efficiency over original SSA. These results establish its potential as a fundamental building block for energy-efficient Spiking Transformers.
Figures & tables
| Method | Type | Unified Form |
|---|---|---|
| SSA [ 67 ] | Attn | |
| SDSA [ 50 ] | Attn | |
| Meta-SDSA [ 49 ] | Attn | |
| Method | Type | Architecture | Time Step | Param (M) | Energy (mJ) | Top-1 Acc (%) |
| ViT [ 37 ] | ANN | ViT-L/16 | 1 | 304.3 | 80.96 | 79.70 |
| PVT [ 39 ] | ANN | PVT-Large | 1 | 61.4 | 45.08 | 81.70 |
| Spikformer [ 67 ] | SNN | Spikformer-8-384 | 4 | 16.8 | 7.73 | 70.24 |
| Spikformer-8-512 | 4 | 29.7 | 11.57 | 73.38 | ||
| Spikingformer [ 62 ] | SNN | Spikingformer-8-512 | 4 | 29.7 | 4.69 | 74.79 |
| Spikingformer-8-768 | 4 | 66.4 | 13.68 | 75.85 |
| Architecture | Param (M) | Step | MIoU (%) |
|---|---|---|---|
| ResNet-18 [ 55 ] | 15.5 | 1 | 32.9 |
| PVT-Tiny [ 39 ] | 17.0 | 1 | 35.7 |
| PVT-Small [ 39 ] | 28.2 | 1 | 39.8 |
| DeepLab-V3 [ 56 ] | 68.1 | 1 | 42.7 |
| SDT-V2 [ 49 ] | 16.5 | 1 | 32.3 |
| SDT-V2 [ 49 ] | 16.5 | 4 | 33.6 |
| Methods | Param. (M) | FELT [ 40 ] | FE108 [ 59 ] | VisEvent [ 41 ] | ||||
|---|---|---|---|---|---|---|---|---|
| AUC(%) | PR(%) | AUC(%) | PR(%) | AUC(%) | PR(%) | |||
| STARK [ 48 ] | 28.23 | 1 | 39.6 | 51.7 | 57.4 | 89.2 | 34.1 | 46.8 |
| ARTrack [ 45 ] | 202.56 | 1 | 39.5 | 49.4 | 56.6 | 88.5 | 33.0 | 43.8 |
| OSTrack 256 [ 52 ] | 92.52 | 1 | 35.9 | 45.5 | 54.6 | 87.1 | 32.7 | 46.4 |
| HIPTrack [ 2 ] | 120.41 | 1 | 38.2 | 48.9 | 50.8 | 81.0 | 32.1 | 45.2 |
| SNNTrack [ 60 ] | 31.40 | 5 | - | - | - | - | 35.4 | 50.4 |
| Hyper-parameter | SDT-V1 | QKformer | SDT-V3 |
|---|---|---|---|
| Timestep | 4 | 4 | 4 |
| Epochs | 300 | 200 | 200 |
| Resolution | 224 | 224 | 224 |
| Batch size | 48 | 100 | 600 |
| Optimizer | AdamW | AdamW | LAMB |
| Base Learning rate | 1.5e-5 | 6e-4 | 6e-4 |
| Test Set Condition | QKFormer Baseline | QKFormer + SCA | (Gain) |
|---|---|---|---|
| Standard Image (Full spectrum) | |||
| Low-Pass Filter (High-freq removed) | |||
| High-Pass Filter (High-freq dominant) |
| Architecture | Params (M) | CIFAR-10 Acc. (%) | CIFAR-100 Acc. (%) | CIFAR-10 DVS Acc. (%) | |
|---|---|---|---|---|---|
| SEW-ResNet [ 7 ] | 1.19 | 4 | - | - | 64.80 |
| 8 | - | - | 70.20 | ||
| SEW-ResNet (Ours) | 1.19 | 4 | - | - | 67.79 ( 2.99) |
| 8 | - | - | 72.66 ( 2.46) | ||
| MS-ResNet-18 [ 14 ] | 11.22 | 4 | 94.40 | 75.06 | - |
| MS-ResNet-18 (Ours) | 11.22 | 4 | 96.37 ( 1.97) | 77.15 ( 2.09) | - |