cs.AISep 28, 2026

QuantaSpike: Short-Window Spike-Driven Quantization for Large Language Models

Authors: Bang Hu, Guowei Zhu, Changze Lv, Xiaoqing Zheng, Fengzhe Zhang, Fan Zhang, Wei Cao

Organizations: School of Computer Science, Fudan University, Shanghai, China

Abstract

Large language models (LLMs) achieve strong performance across many tasks but rely on dense multiply-accumulate (MAC) operations during inference, resulting in high energy cost. Spiking neural networks (SNNs) offer an event-driven alternative in which synaptic integration uses lightweight accumulation. However, spike-driven LLM inference remains difficult because outlier-heavy activations typically require long firing windows or auxiliary non-spiking paths. We propose QuantaSpike, a short-window spike-driven quantization framework for LLMs built around Logarithmic Ternary Integrate-and-Fire (LTIF) neurons. LTIF uses ternary events with power-of-two membrane-response quanta, improving the information represented by each firing step while retaining shift-ACC-compatible computation. QuantaSpike combines this neuron with group-adaptive gain and selective outlier admission: normal values use residual LTIF steps, whereas admitted outliers receive one additional onset spike before entering the same residual dynamics. Across OPT and Llama-2, QuantaSpike achieves state-of-the-art or competitive perplexity and zero-shot accuracy among spike-driven LLM quantization methods. It also transfers to newer dense LLMs, remaining close to the FP16 reference on Llama-3-8B and Qwen3-8B under the same four-step firing window. Analytical linear-energy projections show that QuantaSpike reduces the energy of one linear transformation by about 80.0%80.0\% on OPT models and 67.1%67.1\% on Llama-2 models relative to SpikeQuant, providing an accurate and energy-efficient spike-driven path for LLM inference.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Massive Spikes in LLMs are Bias Vectors: Mechanistic Uncovering and Spike-Free Quantization

    Jun 1, 2026Yung-Chin Chen, Chung Peng Lee, Ze-Wei Liou +1Large Language Model QuantizationPost-Training Quantization

  2. Event-Driven Language Models with Sparse Neural Activity for Neuromorphic Hardware

    Aug 31, 2026Simon Richter, Ruhai Lin, Jason Yik +4Dynamic Sparse AttentionNeuromorphic Computing