cs.AISep 28, 2026

QuantaSpike: Short-Window Spike-Driven Quantization for Large Language Models

Authors: Bang Hu, Guowei Zhu, Changze Lv, Xiaoqing Zheng, Fengzhe Zhang, Fan Zhang, Wei Cao

Organizations: School of Computer Science, Fudan University, Shanghai, China

Abstract

Large language models (LLMs) achieve strong performance across many tasks but rely on dense multiply-accumulate (MAC) operations during inference, resulting in high energy cost. Spiking neural networks (SNNs) offer an event-driven alternative in which synaptic integration uses lightweight accumulation. However, spike-driven LLM inference remains difficult because outlier-heavy activations typically require long firing windows or auxiliary non-spiking paths. We propose QuantaSpike, a short-window spike-driven quantization framework for LLMs built around Logarithmic Ternary Integrate-and-Fire (LTIF) neurons. LTIF uses ternary events with power-of-two membrane-response quanta, improving the information represented by each firing step while retaining shift-ACC-compatible computation. QuantaSpike combines this neuron with group-adaptive gain and selective outlier admission: normal values use residual LTIF steps, whereas admitted outliers receive one additional onset spike before entering the same residual dynamics. Across OPT and Llama-2, QuantaSpike achieves state-of-the-art or competitive perplexity and zero-shot accuracy among spike-driven LLM quantization methods. It also transfers to newer dense LLMs, remaining close to the FP16 reference on Llama-3-8B and Qwen3-8B under the same four-step firing window. Analytical linear-energy projections show that QuantaSpike reduces the energy of one linear transformation by about 80.0%80.0\% on OPT models and 67.1%67.1\% on Llama-2 models relative to SpikeQuant, providing an accurate and energy-efficient spike-driven path for LLM inference.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Jun 1, 2026cs.LG

Massive Spikes in LLMs are Bias Vectors: Mechanistic Uncovering and Spike-Free Quantization

Massive activation spikes in Large Language Models (LLMs) severely degrade quantization by stretching dynamic ranges. While prior hypotheses characterize these as high-level scalar biases, we argue that they are merely the scalar intermediates of rigid, structural vector biases in the spike-carrying tokens. We show that these tokens converge to constant vectors after normalization that drive the attention sink and value-state drain mechanisms. We geometrically substantiate this by analyzing the coordination of projection weights: WKW_K contrastively amplifies the vector, WQW_Q aligns semantic tokens toward it, and WVW_V projects it into the spectral null-space. Furthermore, we reveal that the model actively preserves these structural biases against Rotary Positional Embedding (RoPE) perturbations by localizing them in "zones of rotational stability" utilizing low-frequency bands and coherent channel pairs. Leveraging this, we propose INSERTQUANT, a post-training quantization (PTQ) framework that clamps spikes and restores their function via pre-computed template vectors. This renders activations strictly spike-free, enabling robust low-bit quantization with high fidelity. INSERTQUANT achieves parity with state-of-the-art per-tensor quantization methods on LLMs and uniquely generalizes beyond text to other modalities such as ViTs.
Jun 2, 2026cs.NE

Spike-Aware INT8 Execution for Spiking Language Models on Commodity CPUs

Binary spike activations allow a language-model runtime to read only active weight columns and replace multiplications by weight sums. We implement this execution strategy in C++ for an 874M-parameter spike-gated language model. Sparse projections use column-major INT8 weights, integer accumulation, and one scale application per output channel; dense projections retain row-major access and FP32 activations. In a single-thread comparison using an early checkpoint, INT8 achieves 23.31 tokens/s versus 9.82 for FP32, while reducing weight storage from 3355.2 to 1087.4 MiB. A variant using INT4 on dense projections saves a further 17.4% of storage but reduces decode throughput by 46.6%. On an AMD Ryzen 7 5800X, the final INT8 checkpoint achieves 22.63 tokens/s on one thread and 47.90 on four threads; 512-token prefill reaches 94.68 tokens/s on eight threads. A separate ARM output-head case study records higher trimmed decode-window energy metrics for two candidate-verification configurations. The results characterize how activation-specific layouts and quantized kernels support CPU deployment of a spike-gated language model.
Aug 31, 2026cs.NE

Event-Driven Language Models with Sparse Neural Activity for Neuromorphic Hardware

Inference with transformer-based large language models (LLMs) is often limited by the memory-bound KV cache and quadratic attention cost. State-space models (SSMs) mitigate this through linear attention and fixed-size recurrent states, but their large dense linear projections remain computationally expensive even after quantization. We introduce a method that induces sparse neural activity in heavily quantized linear-attention models with minimal performance loss. Activations below a per-projection trainable threshold (±Δ\pm Δ) are nullified while preserving crucial outliers, achieving comparable performance to dense models with up to 4×\times fewer effective arithmetic operations. Targeting a multi-core, multi-chip neuromorphic platform, where event-driven execution converts unstructured sparsity into throughput at both the compute and communication levels, a capability GPU architectures fundamentally lack, we project up to 37×\times higher throughput and 16×\times lower power versus edge GPU inference of a comparable transformer-based model, and up to 5.4×\times improvements over the non-sparsified baseline. These results position sparse, quantized linear-attention models as a natural fit for deploying LLMs on event-driven multi-core platforms.