Most ANN-to-SNN conversion methods rely on a specific correspondence between the source activation and the spiking neuron dynamics. We propose a finite-state continuous-time Markov chain (CTMC) neuron framework whose stationary spike flux can approximate every continuous nonnegative monotone activation function on a compact interval. For a generalized CTMC family with affine input-dependent transitions, we prove uniform approximation to arbitrary accuracy over this function class and derive an explicit approximation error bound. In practice, two- and three-state CTMCs fit ReLU, sigmoid, softplus, and clipped ReLU on the evaluated input ranges, and we evaluate corresponding MLP conversions for each activation with layerwise rate scaling. Moderate clipping improves the conversion cost-accuracy tradeoff on the MNIST MLP and reduces SynOps by 27% on VGG-11/MNIST at matched ANN-SNN accuracy gap criteria, whereas the trend reverses on VGG-11/CIFAR-10. Mean-field and layerwise diagnostics indicate that finite-window sampling and terminal-layer mismatch are the main residual errors. Overall, our results establish finite-state CTMC neurons as a theoretically grounded framework for activation-flexible ANN-to-SNN conversion beyond fixed activation-neuron correspondences.
Figures & tables
Figure 1: From conventional neurons to a finite-state CTMC neuron. A: Rate-coded ANN-to-SNN conversion at the level of one unit. B: IF and LIF membrane trajectories and their current-to-rate curves; IF aligns with the positive linear ReLU branch, whereas leak introduces a rheobase and curvature. C: The practical three-state neuron has base ( B ), gate ( G ), and refractory ( R ) states. Transitions B→G , G→B , G→R (spike), and R→B occur at rates a(H) , b , c(H) , and d , respectively. The displayed trajectories and parameter sweeps illustrate how leak and reset reshape the stationary firing curve ν(H) .
Figure 2: Stationary-rate fits on the displayed input domains. Solid curves are target activations, dashed curves are fitted CTMC rates ν(H) , and grey curves are residuals (right axis, shared range). The reported mean-squared errors are 1.1×10−3 for sigmoid, 2.4×10−3 for ReLU, 1.1×10−2 for softplus, and 2.7×10−3 for ClipReLU10 . The sigmoid uses the three-state family; the piecewise-linear/softplus examples use the low-state variant used in the implementation. These fits demonstrate finite-range flexibility, not exact global representation.
Setting
Activation
Criterion
SynOps/sample
Spikes/neuron
MNIST MLP
sigmoid
<1%
14.4±0.47 k
0.90
MNIST MLP
softplus
<1%
18.1±2.61 k
1.02
MNIST MLP
ReLU
<1%
13.0±0.58 k
0.87
MNIST MLP
ClipReLU4
<1%
9.1 k
0.55
VGG-11/MNIST
ReLU
<0.5%
2.64 G
12.97
VGG-11/MNIST
ClipReLU6
<0.5%
1.93 G
9.51
Table 1: Highlighted operating points. SynOps are raw spike transmissions per sample. “Criterion” is the ANN–SNN accuracy gap in percentage points.
Figure 3: MNIST MLP activation study ( 784 – 256 – 128 – 10 ; error bars: four-seed standard deviations). A: Accuracy gap versus SynOps for sigmoid (source ANN 98.2%), softplus (98.2%) and ReLU (98.3%); stars mark the minimum-cost configurations below a 1% gap: 14.4±0.47 k SynOps at 0.96±0.04% for sigmoid, 18.1±2.61 k SynOps at 0.91±0.01% for softplus and 13.0±0.58 k at 0.81±0.21% for ReLU. B: Per-neuron spike-count quantiles show a compressed sigmoid distribution and a sparse, heavy-tailed ReLU distribution; mean counts are 0.90 and 0.87. C: Pareto frontiers for sigmoid and a two-unit right shift. Matched-configuration arrows show that suppressing the negative-input tail helps at high budget but can trade away accuracy at low budget.
Figure 4: ClipReLUK ablation on the MNIST MLP, using shared ReLU weights and clip-on-forward evaluation. A: Accuracy-gap/SynOps frontiers. B: Minimum SynOps for an accuracy gap below one percentage point; K=4 is best, and K=1 has no configuration satisfying the criterion. C: Layerwise positive-preactivation quantiles with the clip boundary. D: Fraction exceeding the cap, computed over all units (top) and active units (bottom). Error bars denote standard deviations across four training seeds. Because clipping changes the ANN forward map, each K is evaluated against its own source accuracy.
Figure 5: VGG-11/MNIST: ReLU versus ClipReLU6 . Means over four seeds and eight stochastic trials per seed; error bars are the supplied SEM across runs. A: Pareto frontiers; marker shape denotes τ , and open circles mark the minimum-spike point below a 0.5% gap. B: ANN, mean-field, and SNN accuracies; Cases I/II use (T,τ,r)=(50,0.02,3) and (30,0.02,3) . C: Preactivation quantiles across 14 layers.
Figure 6: VGG-11/CIFAR-10: ReLU versus a 95th-percentile clipped variant. Means over four seeds and four stochastic trials per seed; error bars are the supplied SEM. A: Pareto frontiers and points below a 2% gap. B: ANN, mean-field, and SNN accuracies. C: Preactivation quantiles across 14 layers for Cases I/II, (T,τ,r)=(50,0.02,5) and (40,0.02,4) ; the largest discrepancy is terminal and high-quantile.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Fully connected three-state CTMC fit to the softplus transfer curve. Unlike the restricted low-state CTMC used in the main experiments, this model permits transitions between all pairs of states, with a spike counted on transitions into the designated ”base” state. Top: target rate αsoftplus(H) and the stationary spike-rate prediction over H∈[−26,14] . Bottom: residual between the fitted and target rates. The residual remains within ±0.03 Hz over the evaluated range, with mean-squared error 2.39×10−4 .
Figure 8: Accuracy–cost trade-off for conventional integrate-and-fire (IF) conversion on VGG-11/MNIST, comparing the ReLU (left) and ClipReLU6 (right) source networks. Each point represents one conversion configuration averaged over three seeds. Curves sweep the simulation length T under the indicated normalization and initial-membrane settings v0∈{0,0.5} . The Pareto front shows the lowest attainable gap at each cost, with error bars denoting across-seed standard deviation. The dashed line marks the 0.5% gap criterion, and the star marks the lowest-SynOps configuration satisfying it. The selected ReLU configuration reaches a 0.24% gap at 0.38 G SynOps ( p99.9 , v0=0.5 , T=16 ), while ClipReLU6 reaches a 0.38% gap at 1.04 G SynOps ( v0=0 , T=128 ).
Figure 9: Accuracy–cost trade-off for conventional IF conversion on VGG-11/CIFAR-10 with a ReLU source network, averaged over three seeds. The source ANN accuracy is 91.3% . Results are shown for p100.0 and p99.9 normalization with v0∈{0,0.5} while sweeping the simulation length T . A separate Pareto front is shown for each normalization setting. The dashed line marks the 2% ANN–SNN accuracy-gap criterion, and each star marks the lowest-SynOps configuration satisfying it. Under this criterion, the best p99.9 configuration reaches a 1.64% gap at 0.85 G SynOps ( v0=0.5 , T=48 ), while the best p100.0 configuration reaches a 1.24% gap at 1.09 G SynOps ( v0=0.5 , T=96 ).
Figure 10: Finite-time convergence of empirical firing rates to stationary rates for representative low-state CTMC neurons. Left: a 3-state sigmoid neuron evaluated at several input values H . Right: a 2-state ReLU neuron evaluated at several input values H . Shaded regions indicate variability across stochastic trials, and dashed horizontal lines mark the corresponding stationary firing rates. The results illustrate how finite simulation length contributes to the gap between stationary-rate predictions and actual SNN behavior in different time steps.
Figure 11: Transition-count trade-off for CTMC conversion on VGG-11/CIFAR-10, comparing ReLU (left) and ClipReLU6 (right). Solid lines show the Pareto frontiers, and the minimum-transition configurations satisfying the 2% accuracy-gap criterion: ReLU reaches a 1.93% gap at 6.26×106 transitions/sample, while ClipReLU6 reaches a 1.42% gap at 7.32×106 transitions/sample.
This work proposes a hybrid ANN-SNN pipeline that effectively leverages the rich embeddings of pretrained artificial neural networks (ANNs) to enable high-performance spiking neural networks (SNNs). The architecture couples a pretrained EfficientNet encoder with a CoLaNET spiking classifier. We convert the encoder's activations into spike trains via rate-coding and train the subsequent SNN classifier using local, biologically inspired learning rules, bypassing end-to-end gradient propagation. This approach achieves 99.09% accuracy on a 64-class ImageNet benchmark, demonstrating performance on par with conventional deep networks. The work presents a biologically plausible and efficient framework for adapting powerful pretrained encoders to downstream spiking neural network tasks.
ANN-to-SNN conversion offers a practical, training-free route to spiking large language models. However, current pipelines primarily focus on spike-driven realizations for Transformer linear-algebra operations, while providing limited support for key nonlinear operators. This gap limits compatibility with neuromorphic-style execution constraints, where such nonlinearities typically require division, exponentiation, or norm computations that are not naturally supported by standard leaky integrate-and-fire dynamics. To solve this problem, we propose a plug-and-play framework that implements spike-friendly approximations for Transformer nonlinearities and integrates into existing ANN-to-SNN pipelines. Our method decomposes these nonlinear computations into three recurring primitives -- division, exponentiation, and ℓ2 norms -- and realizes them via population computation using LIF neuron groups, combined with lightweight bit-shift scaling to avoid floating-point arithmetic. By composing these primitives as modular operator blocks, our framework supports common Transformer nonlinearities (e.g., Softmax, SiLU, and normalization) without any fine-tuning. Experiments on a range of LLMs Transformers show that selectively replacing the targeted nonlinear operators incurs less than a 1% accuracy drop across all evaluated tasks.
Xinzhe Yuan, Xiang Peng, Bin Gu +1
IASM, Harbin Institute of Technology · IASM, Harbin Institute of Technology, China · School of Artificial Intelligence, Jilin University +1
Spiking Neural Networks (SNNs) enable event-driven computation with sparse activations, but building multimodal Transformers on SNNs is hindered by unstable training in deep spiking stacks and the mismatch between dense softmax attention and spike-based communication. We propose SMM Transformer, an SNN-based multimodal Transformer framework that combines (i)PLMP, a Parallel LIF with Multistage Learnable Parameters neuron and a tailored P-STBP algorithm for stable deep SNN training, (ii) SMSA, an attention-inspired spike-driven token-mixing module that replaces dense pairwise softmax attention with channel-wise spike co-activation and self-compensation, and (iii)SMoE, a spiking mixture-of-experts module for modality-aware fusion. Across visual and multimodal benchmarks, SMM Transformer achieves competitive accuracy compared to ANN baselines. Under a standard MAC/AC arithmetic model, SMSA reduces the estimated operator-level compute energy of the attention module by up to 97%, while whole-model profiling shows more moderate but consistent efficiency gains.
Xiubo Liang, Jinxing Han, Yuke Li +3
School of Software Technology, Zhejiang University, Ningbo, China · NetEase Yidun AI Lab, Hangzhou, China