Large Language Model-Guided Evolutionary Discovery of Native Neural Architectures for Spiking Sequence Modeling
Organizations: Department of Computer Science, City University of Hong Kong
Abstract
Spiking neural networks (SNNs) offer low-energy sequence modeling through sparse, event-driven computation. However, interactions among spike encoding, neuronal dynamics, and information propagation complicate architecture design. Existing SNN sequence models often adapt artificial neural network (ANN) architectures designed for real-valued activations, potentially underusing spike-based communication and temporal state updates, motivating automated discovery of native SNN architectures. Most evolutionary neural architecture search (ENAS) methods operate within predefined configuration spaces, limiting discovery to mechanisms expressible within those spaces. We introduce OpenArchEvo, which uses large language models (LLMs) to evolve executable architecture code in an open program space under spiking-projection constraints. In this space, code differences need not reflect architectural novelty, while direct performance evaluation requires costly training. We construct a three-view representation spanning code, design rationale, and a behavioral fingerprint to support novelty estimation and performance prediction. The search treats predicted performance and estimated novelty as two objectives, using surrogate predictions to select candidates for expensive training evaluations. With an estimated candidate-training cost of 132 V100 GPU-days, the search uncovers multiple native SNN architectures, exemplified by three designs featuring mechanisms such as spike-activity-dependent control of state updates and residual pathways. The discovered NeuroGate surpasses the ANN DeltaNet on WikiText-103, and the discovered architectures reduce estimated architecture-level arithmetic energy by up to 50.6x (LoopMem) relative to a common dense Transformer (ANN) baseline. All code and all discovered architectures will be made publicly available soon.
Figures & tables
| Model | WT103 | LRA accuracy (%) | Energy reduction | ||||||
|---|---|---|---|---|---|---|---|---|---|
| PPL | ListOps | Text | Retrieval | Image | Pathfinder | Path-X | Avg. | ||
| ANN baselines | |||||||||
| DeltaNet ( Yang et al., 2024b ) | 27.5 | 62.2 | 85.1 | 91.3 | 89.8 | 94.2 | 94.9 | 86.2 | – |
| S4D-Lin † ( Gu et al., 2022 ) | – | 60.5 | 87.0 | 91.0 | 87.9 | 94.0 | 92.8 ∗ | 85.5 | – |
| SNN baselines | |||||||||
| S6-based SNN † | – | 55.7 | 77.6 | 88.5 | 80.1 | 83.4 | – | – | – |
| Variant | WT2 PPL | |
| ① Spike-activity-dependent computation | ||
| HomeoResSSM | ||
| Discovered | 57.5 | – |
| Control driven by continuous inputs | 59.2 | |
| Control removed | 64.4 | |
| ② Spiking-specific design choices | ||
| Adapted ASI-Arch | OpenArchEvo | |
| Best fitness | 0.653 | 0.872 |
| Top-10 mean fitness | 0.570 | 0.714 |
| WT103 PPL | 29.6 | 26.4 |
| ListOps accuracy | 55.3 | 61.8 |
| Text accuracy | 82.3 | 82.4 |
| Retrieval accuracy | 85.3 | 90.1 |
| Selection method | Top-1 PPL | Top-5 PPL | Novelty | Variance |
|---|---|---|---|---|
| Evolution + surrogate | 56.2 | 57.4 | 0.534 | 0.0281 |
| Single island + surrogate | 57.3 | 57.9 | 0.513 | 0.0272 |
| Sampling + surrogate | 57.3 | 57.8 | 0.520 | 0.0314 |
| Sampling only | 57.5 | 58.5 | 0.511 | 0.0414 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Method and source location | Encoded architectural choices | Relation to the discovered computations |
|---|---|---|
| AutoSNN ( Na et al., 2022 ) , Sec. 4.1 | Five block choices in a fixed backbone: skip, spiking convolution, and spiking residual blocks, with specified kernel sizes. | The block menu needs a spike-statistic controller and its target connection to select the activity-dependent update or residual modulation. |
| SNASNet ( Kim et al., 2022 ) , Cell Search Strategy | Four-node cells with zero, skip, convolution, and pooling operations on forward and cross-time backward edges. | Temporal feedback is already permitted. A spike-statistic reduction and learned multiplicative control are additional operations beyond this edge menu. |
| MSE-NAS ( Pan et al., 2025 ) , Sec. III, Fig. 1 | A multiscale genotype selects layer operations, excitatory/inhibitory types, motifs, and global connections under a specified decoder. | These choices alter operations and connectivity; the decoder would need to expose activity-conditioned control equations and their attachment points. |
| EQ-SpikeLM ( Zhang et al., 2026 ) , Sec. IV-B.1, Eqs. (14)–(16) | Per-layer preserved channel ratios for query–key, value, and feed-forward projections in a pretrained spiking language model. | Channel pruning changes widths within the existing computation; it does not introduce a new spike-derived control dependency. |
| ENAS (efficient NAS), recurrent space ( Pham et al., 2018 ) , Secs. 2.1, 3.1 | Predecessor and activation choices in a recurrent cell, with a prescribed highway-gating construction. | Recurrence and multiplicative gating are present. Spike-statistic inputs and revisions to the controller’s equation or target require extending the template. |
| DARTS, recurrent space ( Liu et al., 2019 ) , Sec. 3.1.2 | Operation choices over linear transforms and activations, identity, and zero, within a recurrent-cell template with highway bypasses. | Selecting operations does not itself expose the spike-reduction and control-target relation; these must be added to the operations or template. |
| Group | Measurement | Type | Description |
|---|---|---|---|
| Module statistics | Params | C | Total trainable parameters |
| SpkLinear | C | Spiking linear module count | |
| Conv1d | C | 1-D convolution count | |
| GateRatio | C | Gating layer fraction | |
| Structural statistics | Depth | F | Shape-changing layer count |
| Expand | F | Dim-increase steps |
| Model | Family | Venue | Ref. | Model | Family | Venue | Ref. |
|---|---|---|---|---|---|---|---|
| RetNet | Foundational | arXiv 2023 | ( Sun et al., 2023 ) | S4 | SSM | ICLR 2022 | ( Gu et al., 2022 ) |
| LightNet | Foundational | TMLR 2026 | ( Qin et al., 2026 ) | Mamba (S6) | SSM | COLM 2024 | ( Gu & Dao, 2024 ) |
| GLA | Modern RNN | ICML 2024 | ( Yang et al., 2024a ) | Mamba2 (SSD) | SSM | ICML 2024 | ( Dao & Gu, 2024 ) |
| HGRN | Modern RNN | NeurIPS 2023 | ( Qin et al., 2023 ) | DeltaNet | Delta Rule | NeurIPS 2024 | ( Yang et al., 2024b ) |
| HGRN2 | Modern RNN | COLM 2024 | ( Qin et al., 2024 ) | DeltaFormer | Delta Rule | arXiv 2025 | ( Zhong et al., 2025 ) |
| RWKV-6 | Modern RNN | COLM 2024 | ( Peng et al., 2024a ) | Gated DeltaNet | Delta Rule | ICLR 2025 | ( Yang et al., 2025 ) |
| Setting | Value |
|---|---|
| LLM / maximum generation tokens | Gemini-2.5-Flash / 32,768 |
| Outer iterations / accepted candidates per iteration | 10 / 1,080 |
| Batch limit / evaluation budget | 16 / 160 |
| Islands / retained capacity | 10 / approximately 20 per island |
| Population trimming threshold | Twice the retained capacity |
| Inter-island parent sampling probability | 0.5 |
| Setting | WT2 | ListOps |
|---|---|---|
| Block layers / width | 6 / 256 | 2 / 128 |
| Attention heads | 8 | 4 |
| FFN expansion | 4 | 2 |
| Dropout | 0.2 | 0.1 |
| Batch size | 16 | 32 |
| Epochs | 30 | 25 |
| Architecture | (%) | SOPs (G) | MACs (G) | Energy (mJ) | Reduction |
|---|---|---|---|---|---|
| Dense Transformer | – | – | 9.809 | 45.12 | 1.0 |
| Dyn-SSM | 9.3 | 0.800 | 0.140 | 1.36 | 33.1 |
| SpikingDeltaNet | 9.8 | 0.883 | 0.225 | 1.83 | 24.7 |
| SpikingMamba2 | 12.2 | 1.247 | 0.420 | 3.05 | 14.8 |
| SpikingCondFuse | 8.3 | 0.782 | 0.367 | 2.39 | 18.9 |
| NeuroGate | 4.6 | 0.424 | 0.226 | 1.42 | 31.7 |