QuantaSpike: Short-Window Spike-Driven Quantization for Large Language Models
Organizations: School of Computer Science, Fudan University, Shanghai, China
Abstract
Large language models (LLMs) achieve strong performance across many tasks but rely on dense multiply-accumulate (MAC) operations during inference, resulting in high energy cost. Spiking neural networks (SNNs) offer an event-driven alternative in which synaptic integration uses lightweight accumulation. However, spike-driven LLM inference remains difficult because outlier-heavy activations typically require long firing windows or auxiliary non-spiking paths. We propose QuantaSpike, a short-window spike-driven quantization framework for LLMs built around Logarithmic Ternary Integrate-and-Fire (LTIF) neurons. LTIF uses ternary events with power-of-two membrane-response quanta, improving the information represented by each firing step while retaining shift-ACC-compatible computation. QuantaSpike combines this neuron with group-adaptive gain and selective outlier admission: normal values use residual LTIF steps, whereas admitted outliers receive one additional onset spike before entering the same residual dynamics. Across OPT and Llama-2, QuantaSpike achieves state-of-the-art or competitive perplexity and zero-shot accuracy among spike-driven LLM quantization methods. It also transfers to newer dense LLMs, remaining close to the FP16 reference on Llama-3-8B and Qwen3-8B under the same four-step firing window. Analytical linear-energy projections show that QuantaSpike reduces the energy of one linear transformation by about on OPT models and on Llama-2 models relative to SpikeQuant, providing an accurate and energy-efficient spike-driven path for LLM inference.
Figures & tables
| Metric | OPT-1.3B | OPT-2.7B | Llama-2-7B | Llama-2-13B | ||||||||||||||
| FP16 | SQ | Kirin | QS | FP16 | SQ | Kirin | QS | FP16 | SL | SQ | Kirin | QS | FP16 | SL | SQ | Kirin | QS | |
| Bits | W16 A16 | W4A (4&5) | W4A (4&8) | W4A (4&5) | W16 A16 | W4A (4&5) | W4A (4&8) | W4A (4&5) | W16 A16 | W4A (4&8) | W4A (4&5) | W4A (4&8) | W4A (4&5) | W16 A16 | W4A (4&8) | W4A (4&5) | W4A (4&8) | W4A (4&5) |
| Timestep | – | 32 | 16 | 4 | – | 32 | 16 | 4 | – | 256 | 32 | 16 | 4 | – | 256 | 32 | 16 | 4 |
| WT-2 | 14.63 | 16.38 | 16.33 | 16.38 | 12.47 | 13.15 | 12.89 | 12.76 | 5.68 | 11.36 | 5.80 | 5.71 | 5.78 | 4.88 | 9.71 | 5.04 | 5.03 | 5.11 |
| C4 | 14.72 | 16.56 | 16.26 | 16.34 | 13.17 | 13.91 | 13.72 | 13.79 | 7.08 | 15.87 | 7.33 | 7.25 | 7.23 | 6.46 | 12.10 | 6.65 | 6.61 | 6.61 |
| Model | Method | SNN Conv. | PIQA | ARC- Easy | ARC- Challenge | HellaSwag | WinoGrande | Avg. |
|---|---|---|---|---|---|---|---|---|
| OPT-1.3B | FP16 | ✗ | 72.52 | 50.84 | 29.78 | 53.72 | 59.75 | 53.32 |
| Vanilla Quant. | ✗ | 67.79 | 42.93 | 25.26 | 45.94 | 54.93 | 47.37 | |
| GPTQ | ✗ | 68.34 | 46.17 | 27.65 | 46.53 | 55.26 | 48.79 | |
| SpikeQuant | ✓ | 71.49 | 49.66 | 27.22 | 51.46 | 58.56 | 51.68 | |
| Kirin | ✓ | 71.86 | 48.78 | 28.67 | 52.25 | 57.77 | 51.87 | |
| QuantaSpike | ✓ | 71.84 | 49.23 | 28.79 | 52.34 | 57.70 | 51.98 |
| Method | Llama-3-8B | Qwen3-8B | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| FP16 | Uniform | AWQ | GPTQ | QuantaSpike | FP16 | Uniform | AWQ | GPTQ | QuantaSpike | |
| SNN Conv. | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ |
| WikiText-2 | 6.14 | 6.88 | 6.55 | 7.26 | 6.98 | 9.71 | 10.08 | 10.08 | 10.23 | 10.04 |
| ZS Avg. | 72.73 | 71.34 | 71.33 | 70.14 | 71.25 | 71.58 | 70.01 | 70.66 | 70.59 | 69.41 |
| Ablation | Model | Variant | PPL | Ratio | ||
|---|---|---|---|---|---|---|
| Onset spike | OPT-1.3B | QuantaSpike | 16.38 | 2.90% | 1.34 | 0.262 |
| w/o onset spike | 18.78 | 0.00% | 1.18 | 0.230 | ||
| Llama-2-7B | QuantaSpike | 5.78 | 0.77% | 2.08 | 0.407 | |
| w/o onset spike | 6.04 | 0.00% | 1.95 | 0.382 | ||
| Magnitude gate | Llama-2-7B | QuantaSpike | 5.78 | 0.77% | 2.08 | 0.407 |
| w/o magnitude gate | 39.47 | 4.40% | 2.27 | 0.445 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Model family | Max. window | ||||
|---|---|---|---|---|---|
| OPT | 3 | 64 | 3.0 | 0.5 | 4 |
| Llama-2 | 3 | 64 | 3.0 | 0.5 | 4 |
| Llama-3 | 3 | 32 | 3.0 | 0.5 | 4 |
| Qwen3 dense | 3 | 32 | 3.0 | 0.6 | 4 |
| Qwen3 MoE | 3 | 32 | 3.0 | 0.6 | 4 |
| Model | WikiText-2 PPL | |||
| Llama-3-8B | 3 | 64 | 0.5 | 7.211 |
| 3 | 32 | 0.5 | 6.979 | |
| 3 | 32 | 0.6 | 7.002 | |
| 4 | 32 | 0.5 | 6.958 | |
| Qwen3-8B | 3 | 64 | 0.5 | 10.109 |
| 3 | 32 | 0.5 | 10.046 |
| Model | Method | PPL | PIQA | ARC-Easy | ARC-Challenge | HellaSwag | WinoGrande | Avg. |
|---|---|---|---|---|---|---|---|---|
| Llama-3-8B | FP16 | 6.14 | 80.85 | 77.69 | 53.33 | 79.15 | 72.61 | 72.73 |
| Uniform | 6.88 | 79.49 | 77.15 | 50.43 | 77.79 | 71.82 | 71.34 | |
| QuantaSpike | 6.98 | 80.15 | 77.05 | 50.11 | 77.47 | 71.47 | 71.25 | |
| Qwen3-8B | FP16 | 9.71 | 77.75 | 80.93 | 56.57 | 74.93 | 67.72 | 71.58 |
| Uniform | 10.08 | 75.90 | 76.14 | 53.41 | 73.72 | 65.90 | 69.01 | |
| QuantaSpike | 10.04 | 76.42 | 76.55 | 54.72 | 73.19 | 66.18 | 69.41 |
| Model | Accounting | ||||
|---|---|---|---|---|---|
| OPT-1.3B | All residual steps counted | 2.90% | 0.00% | 3.03 | 0.593 |
| QuantaSpike (nonzero only) | 2.90% | 56.37% | 1.34 | 0.262 | |
| Llama-2-7B | All residual steps counted | 0.77% | 0.00% | 3.01 | 0.588 |
| QuantaSpike (nonzero only) | 0.77% | 30.92% | 2.08 | 0.407 |
| Method | OPT-1.3B | OPT-2.7B | Llama-2-7B | Llama-2-13B |
|---|---|---|---|---|
| Quantization | 4.732 | 7.393 | – | – |
| OneBit | 13.237 | 20.683 | – | – |
| LRQuant | 13.690 | 21.390 | – | – |
| GPTQ | 14.227 | 22.229 | – | – |
| SmoothQuant | – | – | 18.927 | 29.573 |
| OmniQuant | – | – | 18.927 | 29.573 |
| Model | Variant | PPL | Outlier MSE | Ratio | ||
|---|---|---|---|---|---|---|
| OPT-1.3B | QuantaSpike | 16.380 | 0.00326 | 2.90% | 1.337 | 0.262 |
| w/o onset | 18.784 | 0.05322 | 0.00% | 1.178 | 0.230 | |
| Llama-2-7B | QuantaSpike | 5.779 | 0.00334 | 0.77% | 2.082 | 0.407 |
| w/o onset | 6.039 | 0.02093 | 0.00% | 1.950 | 0.382 |
| Variant | PPL | Ratio | ||
|---|---|---|---|---|
| QuantaSpike | 5.779 | 0.77% | 2.082 | 0.407 |
| w/o magnitude gate | 39.470 | 4.40% | 2.270 | 0.445 |
| Variant | Quantized scope | Calibrated | Skipped | PPL |
|---|---|---|---|---|
| FP16 | None | – | – | 11.970 |
| Attention | Attn. + shared dense | 192 | 0 | 11.900 |
| Experts | Routed experts | 10,578 | 7,854 | 12.254 |
| Experts + diverse calib. | Routed experts | 13,488 | 4,944 | 12.254 |