Multi-Depth Temporal Fusion for Feedforward, Locally Trained Spiking Neural Networks
Organizations: Department of Information Engineering, University of Padua Padua, Italy
Abstract
We propose a new spiking neural network (SNN) design to process static images and event streams using time-to-first-spike (TTFS) latencies. Our key research question is which architectural choices best accommodate local and online learning in multi-layer convolutional SNNs. This question is addressed via an original framework combining residual-like connections with multi-depth feature aggregation and consensus. The full SNN pipeline features an early-vision front end, to convert raw visual data into sparse spike latencies, a four-layer convolutional backbone trained layerwise with unsupervised spike-timing-dependent plasticity (STDP), a deterministic Multi-Depth Temporal Fusion (MDTF) and a final classifier trained with reward-modulated spike-timing-dependent plasticity (R-STDP). Rather than replacing early features in deeper layers, the proposed MDTF preserves early temporal evidence, adding sparse residual events from intermediate layers, and incorporating deeper features only when they agree in time with earlier representations. The resulting architecture is experimentally validated across MNIST, Fashion-MNIST, CIFAR-10, and N-MNIST, delivering strong classification performance under a fully local learning regime. Selective multi-depth fusion significantly outperforms traditional STDP/R-STDP baselines on higher-variability visual tasks (achieving +18.2 pp on Fashion-MNIST and +29.2 pp on CIFAR-10). Furthermore, activity-budget analyses show that the network retains high accuracy even when removing a large fraction of late or weak spike events, confirming its high data efficiency and reduced event-processing requirements. The codebase is publicly available at github.com/aidinattar/multi-depth-temporal-fusion-snn.
Figures & tables
| Stage | Plasticity | Role in the pipeline |
|---|---|---|
| Visual front end | None; deterministic preprocessing | Converts static images or event streams into label-free latency representations. Static-input statistics are estimated once from the training split, whereas event-based normalization is applied independently to each sample. |
| Early convolutional stages | Unsupervised STDP in - ; deterministic - export | To learn early temporal features and produce the preserved representation and the intermediate representation ; each trained spiking stage is frozen before subsequent stages are fitted. |
| Deep convolutional path | Unsupervised STDP in - | Further processes to produce the deep representation , used together with by the temporal agreement mechanism. |
| Sparse Multi-Depth Temporal Fusion | None; deterministic selection and agreement | Preserves , extracts the sparse residual contribution from , and adds the agreement contribution obtained from temporally consistent events in and . |
| Multi-prototype readout | Reward-modulated STDP | Learns class decisions from the frozen fused representation using target-side reinforcement, hard-negative anti-STDP, and multiple prototypes per class. |
| Dataset | Input structure | Code dim. | Events/sample | Density | Accuracy (%) |
|---|---|---|---|---|---|
| MNIST | grayscale digits | ||||
| Fashion-MNIST | grayscale apparel | ||||
| CIFAR-10 | RGB natural images | ||||
| N-MNIST | polarity event streams |
| Dataset | Front-end configuration | Main configuration change | Accuracy |
|---|---|---|---|
| MNIST | simple latency | local decorrelation and polarity processing | |
| MNIST | full without signed context | signed-context gate | |
| MNIST | full without polarity balancing | local polarity-balance gates | |
| MNIST | full static front end | - | |
| Fashion-MNIST | simple latency | local decorrelation and polarity processing | |
| Fashion-MNIST | full without signed context | signed-context gate |
| Dataset | Local STDP/R-STDP baseline | Proposed full model | pp | |
|---|---|---|---|---|
| MNIST | ||||
| Fashion-MNIST | ||||
| CIFAR-10 | ||||
| N-MNIST |
| Dataset | 100 | 500 | 1k | 5k | 10k | 20k | 30k |
|---|---|---|---|---|---|---|---|
| MNIST | |||||||
| Fashion-MNIST | |||||||
| CIFAR-10 | |||||||
| N-MNIST |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Parameter | MNIST/Fashion-MNIST | CIFAR-10 | N-MNIST |
|---|---|---|---|---|
| Input source | Raw data | grayscale images | RGB images | polarity event stream |
| Latency encoder | Representation | calibrated local-decorrelation polarity TTFS | calibrated local-decorrelation polarity TTFS | event-map TTFS |
| Latency encoder | Output channels | 2 | 6 | 10 |
| Local decorrelation | fit-patch , stride, | , , | , , | – |
| Local decorrelation | PCA component fraction, max. patches, fit batch | , , | , , | – |
| Polarity calibration | split and scaling | positive/negative; training-set min–max per channel and location | positive/negative; training-set min–max per channel and location | native ON/OFF channels |
| Block | Setting | Input | Output | or op | or | Top- | Ep. | ||
|---|---|---|---|---|---|---|---|---|---|
| MNIST/Fashion-MNIST | ; pool | 5.0 | 0 | - | 8 | ||||
| CIFAR-10 | ; pool | 10.0 | 0 | - | 18 | ||||
| N-MNIST | ; pool | 4.0 | 0 | - | 2 | ||||
| MNIST/Fashion-MNIST | ; pool | 24.0 | 32 | - | 3 | ||||
| CIFAR-10 | ; pool | 24.0 | 32 | - | 3 | ||||
| N-MNIST | ; pool | 8.0 | 32 | - | 2 |
| Dataset | Proto./class | Updated proto. | Non-target scale | Correct scale | |||||
|---|---|---|---|---|---|---|---|---|---|
| MNIST | 80 | 8 | 120 | 7 | 0.25 | 0.0035 | 0.0075 | ||
| Fashion-MNIST | 40 | 4 | 223 | 5 | 0.32 | 0.0050 | 0.1000 | ||
| CIFAR-10 | 40 | 4 | 400 | 5 | 0.32 | 0.0050 | 0.1000 | ||
| N-MNIST | 60 | 6 | 243 | 9 | 0.35 | 0.0030 | 0.0050 |