cs.LGOct 7, 2026

Rethinking the Tradeoff Between Temporal Encoding and Nonlinear Computation in Spiking Language Models

Authors: Hanfei Liu, Shuchang Feng, Yanxia Chen, Changzeng Fu, Shiqi Zhao

Organizations: Northeastern University, 110819, Shenyang, China · Hebei Key Laboratory of Marine Perception Network and Data Processing, 066004, Qinhuangdao, China

Abstract

Spiking language models face a tradeoff between representing continuous semantic features over short temporal windows and retaining costly nonlinear attention operations. We introduce Spora, which jointly designs spike encodings and attention operators. Binary temporal weights let TT spikes represent compositional values with up to TT bits of capacity, compared with O(log⁡2T)O(\log_2 T) bits for spike-count readout. Unipolar Binary Spiking (UBS) uses thresholds and spike-triggered residual decay to produce non-negative integer codes; Bipolar Binary Spiking (BBS) separates sign and magnitude and learns a scale for signed activations. These representations support accumulation-and-shift dot products and integer-exponent mappings in attention. With four time steps, Spora achieves 76.6 average GLUE score and 44.1 CoLA MCC, improving over SpikeLM by 1.2 and 6.2 points, respectively. Extending BBS to six steps raises these scores to 78.2 and 47.4. Conditional-decay analysis, matched-budget activation-quantization comparisons, event-workload statistics, and fixed-point evaluation further characterize the connection between encoding fidelity and computational cost.

Figures & tables

Explore similar work

CardsList
  1. QuantaSpike: Short-Window Spike-Driven Quantization for Large Language Models

    Sep 28, 2026Bang Hu, Guowei Zhu, Changze Lv +4LLM QuantizationSpiking Neural Networks