Large language models (LLMs) achieve strong performance across many tasks but rely on dense multiply-accumulate (MAC) operations during inference, resulting in high energy cost. Spiking neural networks (SNNs) offer an event-driven alternative in which synaptic integration uses lightweight accumulation. However, spike-driven LLM inference remains difficult because outlier-heavy activations typically require long firing windows or auxiliary non-spiking paths. We propose QuantaSpike, a short-window spike-driven quantization framework for LLMs built around Logarithmic Ternary Integrate-and-Fire (LTIF) neurons. LTIF uses ternary events with power-of-two membrane-response quanta, improving the information represented by each firing step while retaining shift-ACC-compatible computation. QuantaSpike combines this neuron with group-adaptive gain and selective outlier admission: normal values use residual LTIF steps, whereas admitted outliers receive one additional onset spike before entering the same residual dynamics. Across OPT and Llama-2, QuantaSpike achieves state-of-the-art or competitive perplexity and zero-shot accuracy among spike-driven LLM quantization methods. It also transfers to newer dense LLMs, remaining close to the FP16 reference on Llama-3-8B and Qwen3-8B under the same four-step firing window. Analytical linear-energy projections show that QuantaSpike reduces the energy of one linear transformation by about 80.0% on OPT models and 67.1% on Llama-2 models relative to SpikeQuant, providing an accurate and energy-efficient spike-driven path for LLM inference.
Figures & tables
Figure 1: Zero-shot accuracy versus linear energy on Llama-2-7B.
Figure 2: Overview of QuantaSpike. QuantaSpike maps LLM activations into the LTIF membrane domain, admits only necessary outlier states to an adaptive onset spike, and represents activations with short-window ternary LTIF spikes. These spikes are processed via sign selection, power-of-two shifts, and accumulation, ensuring a purely spike-driven linear path free of MAC operations.
Metric
OPT-1.3B
OPT-2.7B
Llama-2-7B
Llama-2-13B
FP16
SQ
Kirin
QS
FP16
SQ
Kirin
QS
FP16
SL
SQ
Kirin
QS
FP16
SL
SQ
Kirin
QS
Bits
W16 A16
W4A (4&5)
W4A (4&8)
W4A (4&5)
W16 A16
W4A (4&5)
W4A (4&8)
W4A (4&5)
W16 A16
W4A (4&8)
W4A (4&5)
W4A (4&8)
W4A (4&5)
W16 A16
W4A (4&8)
W4A (4&5)
W4A (4&8)
W4A (4&5)
Timestep
–
32
16
4
–
32
16
4
–
256
32
16
4
–
256
32
16
4
WT-2 ↓
14.63
16.38
16.33
16.38
12.47
13.15
12.89
12.76
5.68
11.36
5.80
5.71
5.78
4.88
9.71
5.04
5.03
5.11
C4 ↓
14.72
16.56
16.26
16.34
13.17
13.91
13.72
13.79
7.08
15.87
7.33
7.25
7.23
6.46
12.10
6.65
6.61
6.61
Table 1: Perplexity on WikiText-2(WT-2) and C4 across model scales. We compare SpikeQuant (SQ), SpikeLLM (SL) and Kirin with our method QuantaSpike(QS)
Model
Method
SNN Conv.
PIQA
ARC- Easy
ARC- Challenge
HellaSwag
WinoGrande
Avg.
OPT-1.3B
FP16
✗
72.52
50.84
29.78
53.72
59.75
53.32
Vanilla Quant.
✗
67.79
42.93
25.26
45.94
54.93
47.37
GPTQ
✗
68.34
46.17
27.65
46.53
55.26
48.79
SpikeQuant
✓
71.49
49.66
27.22
51.46
58.56
51.68
Kirin
✓
71.86
48.78
28.67
52.25
57.77
51.87
QuantaSpike
✓
71.84
49.23
28.79
52.34
57.70
51.98
Table 2: Zero-shot accuracy on OPT and Llama-2 models. FP16 and QuantaSpike are evaluated in our pipeline; other baselines are reported from prior papers. Best and second-best results among quantized or spike-driven methods are bolded and underlined.
Method
Llama-3-8B
Qwen3-8B
FP16
Uniform
AWQ
GPTQ
QuantaSpike
FP16
Uniform
AWQ
GPTQ
QuantaSpike
SNN Conv.
✗
✗
✗
✗
✓
✗
✗
✗
✗
✓
WikiText-2 ↓
6.14
6.88
6.55
7.26
6.98
9.71
10.08
10.08
10.23
10.04
ZS Avg. ↑
72.73
71.34
71.33
70.14
71.25
71.58
70.01
70.66
70.59
69.41
Table 3: WikiText-2 perplexity and zero-shot accuracy on newer dense LLMs. Comparison of FP16, Uniform, AWQ, GPTQ and our QuantaSpike.
Figure 3: Estimated energy of one linear transformation under the operation-cost model adopted by SpikeQuant and Kirin.
Ablation
Model
Variant
PPL
po
mˉ
Ratio
Onset spike
OPT-1.3B
QuantaSpike
16.38
2.90%
1.34
0.262
w/o onset spike
18.78
0.00%
1.18
0.230
Llama-2-7B
QuantaSpike
5.78
0.77%
2.08
0.407
w/o onset spike
6.04
0.00%
1.95
0.382
Magnitude gate
Llama-2-7B
QuantaSpike
5.78
0.77%
2.08
0.407
w/o magnitude gate
39.47
4.40%
2.27
0.445
Table 4: Ablation studies. The onset spike improves outlier reconstruction, while the magnitude gate controls false outlier admission.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Model family
n
G
λ
γ
Max. window
OPT
3
64
3.0
0.5
4
Llama-2
3
64
3.0
0.5
4
Llama-3
3
32
3.0
0.5
4
Qwen3 dense
3
32
3.0
0.6
4
Qwen3 MoE
3
32
3.0
0.6
4
Appendix
Table 5: Model-specific QuantaSpike configurations. n is the number of residual LTIF firing steps, G is the group size, λ is the MAD coefficient, and γ is the magnitude-gate coefficient.
Model
n
G
γ
WikiText-2 PPL ↓
Llama-3-8B
3
64
0.5
7.211
3
32
0.5
6.979
3
32
0.6
7.002
4
32
0.5
6.958
Qwen3-8B
3
64
0.5
10.109
3
32
0.5
10.046
Appendix
Table 6: Hyperparameter sensitivity on newer dense LLMs. The selected four-step configurations are marked in bold.
Model
Method
PPL ↓
PIQA
ARC-Easy
ARC-Challenge
HellaSwag
WinoGrande
Avg. ↑
Llama-3-8B
FP16
6.14
80.85
77.69
53.33
79.15
72.61
72.73
Uniform
6.88
79.49
77.15
50.43
77.79
71.82
71.34
QuantaSpike
6.98
80.15
77.05
50.11
77.47
71.47
71.25
Qwen3-8B
FP16
9.71
77.75
80.93
56.57
74.93
67.72
71.58
Uniform
10.08
75.90
76.14
53.41
73.72
65.90
69.01
QuantaSpike
10.04
76.42
76.55
54.72
73.19
66.18
69.41
Appendix
Table 7: Full results on newer dense LLMs. We report WikiText-2 perplexity and zero-shot accuracy. Normalized accuracy is used for PIQA, ARC-Easy, ARC-Challenge, and HellaSwag; accuracy is used for WinoGrande. All three methods are evaluated in the same pipeline.
Model
Accounting
po
pz
mˉ
Ract
OPT-1.3B
All residual steps counted
2.90%
0.00%
3.03
0.593
QuantaSpike (nonzero only)
2.90%
56.37%
1.34
0.262
Llama-2-7B
All residual steps counted
0.77%
0.00%
3.01
0.588
QuantaSpike (nonzero only)
0.77%
30.92%
2.08
0.407
Appendix
Table 8: Effect of zero-spike silence on event accounting. All residual steps counted is an accounting reference that treats silent residual steps as active; it preserves the same representation and perplexity as QuantaSpike. po is the onset-event rate, pz is the silent residual-step rate, mˉ is the average counted events per activation, and Ract is the normalized accumulation-cost ratio.
Method
OPT-1.3B
OPT-2.7B
Llama-2-7B
Llama-2-13B
Quantization
4.732
7.393
–
–
OneBit
13.237
20.683
–
–
LRQuant
13.690
21.390
–
–
GPTQ
14.227
22.229
–
–
SmoothQuant
–
–
18.927
29.573
OmniQuant
–
–
18.927
29.573
Appendix
Table 9: Linear-energy values used in Fig. 3 ( μ J). Prior-method values follow the corresponding reported linear-layer analysis. QuantaSpike values are obtained from the operation-level projection above. ∗ indicates an estimate using same-family firing statistics. Lower is better.
Model
Variant
PPL
Outlier MSE
po
mˉ
Ratio
OPT-1.3B
QuantaSpike
16.380
0.00326
2.90%
1.337
0.262
w/o onset
18.784
0.05322
0.00%
1.178
0.230
Llama-2-7B
QuantaSpike
5.779
0.00334
0.77%
2.082
0.407
w/o onset
6.039
0.02093
0.00%
1.950
0.382
Appendix
Table 10: Ablation of the adaptive onset spike. Removing the onset response saves a small event budget but sharply increases outlier reconstruction error.
Variant
PPL
po
mˉ
Ratio
QuantaSpike
5.779
0.77%
2.082
0.407
w/o magnitude gate
39.470
4.40%
2.270
0.445
Appendix
Table 11: Ablation of the magnitude gate on Llama-2-7B. MAD-only admission spends more onset events but severely degrades reconstruction.
Variant
Quantized scope
Calibrated
Skipped
PPL
FP16
None
–
–
11.970
Attention
Attn. + shared dense
192
0
11.900
Experts
Routed experts
10,578
7,854
12.254
Experts + diverse calib.
Routed experts
13,488
4,944
12.254
Appendix
Table 12: Diagnostic MoE results on Qwen3-30B-A3B. The router remains in high precision. Evaluation uses eight WikiText-2 sequences of length 512 ; lower perplexity is better.
Massive activation spikes in Large Language Models (LLMs) severely degrade quantization by stretching dynamic ranges. While prior hypotheses characterize these as high-level scalar biases, we argue that they are merely the scalar intermediates of rigid, structural vector biases in the spike-carrying tokens. We show that these tokens converge to constant vectors after normalization that drive the attention sink and value-state drain mechanisms. We geometrically substantiate this by analyzing the coordination of projection weights: WK contrastively amplifies the vector, WQ aligns semantic tokens toward it, and WV projects it into the spectral null-space. Furthermore, we reveal that the model actively preserves these structural biases against Rotary Positional Embedding (RoPE) perturbations by localizing them in "zones of rotational stability" utilizing low-frequency bands and coherent channel pairs. Leveraging this, we propose INSERTQUANT, a post-training quantization (PTQ) framework that clamps spikes and restores their function via pre-computed template vectors. This renders activations strictly spike-free, enabling robust low-bit quantization with high fidelity. INSERTQUANT achieves parity with state-of-the-art per-tensor quantization methods on LLMs and uniquely generalizes beyond text to other modalities such as ViTs.
Yung-Chin Chen, Chung Peng Lee, Ze-Wei Liou +1
Princeton University, NJ, USA · EnCharge AI, CA, USA
Binary spike activations allow a language-model runtime to read only active weight columns and replace multiplications by weight sums. We implement this execution strategy in C++ for an 874M-parameter spike-gated language model. Sparse projections use column-major INT8 weights, integer accumulation, and one scale application per output channel; dense projections retain row-major access and FP32 activations. In a single-thread comparison using an early checkpoint, INT8 achieves 23.31 tokens/s versus 9.82 for FP32, while reducing weight storage from 3355.2 to 1087.4 MiB. A variant using INT4 on dense projections saves a further 17.4% of storage but reduces decode throughput by 46.6%. On an AMD Ryzen 7 5800X, the final INT8 checkpoint achieves 22.63 tokens/s on one thread and 47.90 on four threads; 512-token prefill reaches 94.68 tokens/s on eight threads. A separate ARM output-head case study records higher trimmed decode-window energy metrics for two candidate-verification configurations. The results characterize how activation-specific layouts and quantized kernels support CPU deployment of a spike-gated language model.
Inference with transformer-based large language models (LLMs) is often limited by the memory-bound KV cache and quadratic attention cost. State-space models (SSMs) mitigate this through linear attention and fixed-size recurrent states, but their large dense linear projections remain computationally expensive even after quantization. We introduce a method that induces sparse neural activity in heavily quantized linear-attention models with minimal performance loss. Activations below a per-projection trainable threshold (±Δ) are nullified while preserving crucial outliers, achieving comparable performance to dense models with up to 4× fewer effective arithmetic operations. Targeting a multi-core, multi-chip neuromorphic platform, where event-driven execution converts unstructured sparsity into throughput at both the compute and communication levels, a capability GPU architectures fundamentally lack, we project up to 37× higher throughput and 16× lower power versus edge GPU inference of a comparable transformer-based model, and up to 5.4× improvements over the non-sparsified baseline. These results position sparse, quantized linear-attention models as a natural fit for deploying LLMs on event-driven multi-core platforms.
Simon Richter, Ruhai Lin, Jason Yik +4
Department of Electrical and Computer Engineering, Aarhus University, Aarhus, Denmark · Department of Computer Science and Engineering, University of California, Santa Cruz, Santa Cruz, USA · School of Engineering and Applied Sciences, Harvard University, Cambridge, USA +1