Organizations: School of Life Sciences, Tsinghua University, Beijing, China · School of Future Information Innovation, Fudan University, Shanghai, China
Two-photon calcium imaging is a standard tool for recording large neural populations in vivo, yet inferring spikes accurately across the growing diversity of calcium indicators remains an open problem. Existing supervised methods achieve reasonable in-domain accuracy but generalize poorly to unseen indicators, because different indicators induce distinct fluorescence kinetics and signal statistics while existing architectures remain relatively simple generic temporal regressors without dynamics-matched inductive bias. We propose SpikeSSL, a universal spike inference framework whose temporal backbone is a bank of bidirectional IIR state-space layers broadly motivated by calcium dynamics. A multi-modal conditioning encoder maps indicator identity, sampling rate, and trace-level signal statistics into a global conditioning vector that modulates the backbone via Adaptive Layer Normalization, while a heteroscedastic variance head provides calibrated per-frame uncertainty. On a benchmark with five fixed evaluation splits built from 33 public ground-truth datasets, SpikeSSL achieves state-of-the-art performance in both in-domain and zero-shot leave-one-indicator-out settings. We also develop a biophysical simulation pipeline capable of generating paired fluorescence-spike traces with systematically varied kinetic parameters, spike statistics, response nonlinearities, baseline drift, and noise. Using this pipeline, we synthesize approximately 11,000 simulated traces. Augmenting training with these data effectively closes the cross-indicator domain gap and improves zero-shot generalization. Code is publicly available at https://github.com/detimage123/SpikeSSL.
Figures & tables
Figure 1: Indicator diversity and coverage gaps. (a) Representative fluorescence and spike snippets show markedly different response shapes across calcium indicators and among related GCaMP variants. (b) Indicator medians separate in kinetic and signal-statistic space, while simulated traces populate sparse regions between real datasets. (c) A t-SNE projection of per-window features shows partially overlapping but systematically displaced indicator domains.
Figure 2: Architecture of SpikeSSL. A fluorescence trace is projected by a pointwise convolution and processed by a stack of SpikeSSL blocks. Each block contains a bidirectional IIR scan with learnable decay rates that capture the exponential decay dynamics of calcium indicators, modulated by AdaLN conditioning on indicator identity, sampling rate, and trace-level signal statistics, followed by a Temporal Mixer for local temporal mixing and a SwiGLU FFN for channel mixing. A dilated refinement module aggregates multi-scale temporal context. Three parallel output heads produce a spike rate estimate, a binary spiking gate, and a per-frame uncertainty estimate.
Figure 3: Hierarchical Workflow of the Calcium Data Simulator.
Method
V1-OGB
V1-GCaMP6s
V1-GCaMP8
Other-RCaMP
SC-GCaMP6s
Mean
Params
Corr ↑
Err ↓
Bias ↓
Corr ↑
Err ↓
Bias ↓
Corr ↑
Err ↓
Bias ↓
Corr ↑
Err ↓
Bias ↓
Corr ↑
Err ↓
Bias ↓
Corr ↑
OASIS
0.298
1.157
− 0.481
0.332
0.973
− 0.883
0.493
0.878
− 0.659
0.504
1.054
− 0.508
0.365
0.975
− 0.964
0.398
—
MLSpike
0.363
4.379
− 3.244
0.224
1.017
− 0.928
0.558
0.935
− 0.887
0.354
1.369
− 0.761
0.307
1.005
− 0.919
0.361
—
CASCADE [ 16 ]
0.613
1.117
− 0.250
0.361
1.718
− 0.520
0.825
3.246
− 3.108
0.675
1.060
− 0.388
0.149
1.005
− 0.639
0.524
34K
ENS 2 [ 28 ]
0.641
1.032
− 0.368
0.358
2.265
− 1.354
0.775
1.827
− 1.596
0.695
0.884
− 0.509
0.256
1.116
− 0.154
0.545
145K
SpikeSSL
0.743
0.810
− 0.067
0.579
0.945
− 0.741
0.915
0.565
− 0.017
0.731
1.269
− 0.234
0.736
0.764
− 0.704
0.741
3.9M
Table 1: In-domain resluts. All methods are trained and evaluated on the full ground-truth database with per-dataset train and test splits. Bold / Underline indicate the 1st / 2nd best results.
Method
V1-OGB
V1-GCaMP6s
V1-GCaMP8
Other-RCaMP
SC-GCaMP6s
Mean
Corr ↑
Err ↓
Bias ↓
Corr ↑
Err ↓
Bias ↓
Corr ↑
Err ↓
Bias ↓
Corr ↑
Err ↓
Bias ↓
Corr ↑
Err ↓
Bias ↓
Corr ↑
OASIS
0.298
1.157
− 0.481
0.332
0.973
− 0.883
0.493
0.878
− 0.659
0.504
1.054
− 0.508
0.365
0.975
− 0.964
0.398
MLSpike
0.363
4.379
− 3.244
0.224
1.017
− 0.928
0.558
0.935
− 0.887
0.354
1.369
− 0.761
0.307
1.005
− 0.919
0.361
CASCADE
0.560
1.208
− 0.144
0.304
2.571
1.614
0.757
3.723
3.602
0.551
1.086
− 0.300
0.126
1.083
− 0.442
0.460
ENS 2
0.574
1.005
− 0.366
0.310
3.401
2.654
0.716
2.038
1.715
0.635
0.966
− 0.454
0.178
1.092
− 0.391
0.483
SpikeSSL
0.650
1.979
− 0.675
0.509
0.944
− 0.677
0.853
0.513
− 0.091
0.717
1.172
− 0.606
0.542
0.911
− 0.892
0.654
Table 2: Zero-shot LOIO results. Each column group holds out one of the five evaluation splits during training and evaluates on it without target-domain data. Bold / Underline indicate the 1st / 2nd best results.
Method
V1-OGB
V1-GCaMP6s
V1-GCaMP8
Other-RCaMP
SC-GCaMP6s
Mean
Corr ↑
Err ↓
Bias ↓
Corr ↑
Err ↓
Bias ↓
Corr ↑
Err ↓
Bias ↓
Corr ↑
Err ↓
Bias ↓
Corr ↑
Err ↓
Bias ↓
Corr ↑
CASCADE + Sim.
0.552
2.418
− 1.423
0.439
1.522
− 0.691
0.716
1.408
− 1.018
0.642
1.421
− 0.614
0.405
3.206
− 2.084
0.551
ENS 2 + Sim.
0.486
1.634
− 0.592
0.447
2.146
− 1.418
0.796
1.882
− 1.487
0.588
1.764
− 0.756
0.461
2.291
− 1.054
0.556
SpikeSSL + Sim.
0.697
1.441
− 0.419
0.533
0.931
− 0.697
0.902
0.523
− 0.175
0.708
0.849
− 0.505
0.678
0.771
− 0.653
0.704
Table 3: LOIO results with simulation augmentation (“+ Sim.”). Training is augmented with biophysically simulated calcium traces. Bold / Underline indicate the 1st / 2nd best results.
Figure 4: Qualitative LOIO comparison on GCaMP8s, GCaMP6s, and OGB-1 test neurons. Each column shows a different method, and the grey curve denotes the ground-truth smoothed spike rate.
Figure 8Figure 9
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Conditioning encoder architecture. Three modality-specific branches independently encode their inputs before fusing them into a common representation. The indicator embedding branch maps discrete indicator identity to a 64-dimensional learned vector, where ID 0 is reserved for unknown or unseen indicators. The sampling-rate branch encodes logfs through a two-layer MLP. The signal-statistics branch encodes the 7-dimensional descriptor comprising standard deviation, skewness, excess kurtosis, and autocorrelation at lags 1, 3, 5, and 10 through a two-layer MLP. The three 64-dimensional outputs are concatenated and fused by a 192→64→64 MLP to produce the global conditioning vector c∈R64 .
Figure 8: (a) SpikeSSL block sub-modules. The Temporal Mixer applies a depthwise 1D convolution (kernel width 7) followed by a pointwise convolution, GELU activation, and a residual connection, providing local temporal mixing without indicator-conditioned normalization. The SwiGLU FFN projects to twice the channel width, splits into gate and value tensors, computes SiLU(gate)⊙value , and projects back to model dimension D , followed by a residual connection. (b) Dilated refinement module comprising three residual layers with dilation factors 1, 2, and 4 that successively expand the temporal receptive field at full resolution, each containing a kernel-3 1D convolution, batch normalization, GELU, and dropout.
Figure 9: Representative transfer functions for five commonly used indicators. The shaded region denotes inter-neuron variability, and the dashed diagonal indicates the identity mapping. Indicators are selected because they span distinct kinetic regimes.
Figure 10: LOIO spike inference on a representative OGB-1 neuron (chemical dye with slow calcium decay kinetics). The top row shows the recorded ΔF/F trace. Subsequent rows display spike-rate estimates produced by OASIS, MLSpike, CASCADE, ENS 2 , SpikeSSL, and SpikeSSL+Sim., each overlaid with the smoothed ground-truth spike rate (grey). The Pearson correlation coefficient r is computed over the full recording duration.
Figure 11: LOIO spike inference on a representative GCaMP5k neuron (an early-generation fast green GECI). Panel layout follows Figure 10 .
Figure 12: LOIO spike inference on a representative GCaMP7f neuron (an improved fast green GECI with enhanced sensitivity). Panel layout follows Figure 10 .
Figure 13: LOIO spike inference on a representative GCaMP6s neuron (a slow GECI recorded in a low-activity regime). Panel layout follows Figure 10 .
Figure 14: LOIO spike inference on a representative jGECO1a neuron (a red-shifted GECI with distinct spectral and kinetic properties). Panel layout follows Figure 10 .
Figure 15: LOIO spike inference on a representative GCaMP8f neuron (a next-generation fast GECI with high signal-to-noise ratio). Panel layout follows Figure 10 .
Large-scale calcium imaging has created an opportunity to build foundation-style models for neural population dynamics, but a central question remains unresolved: \textbf{whether a model pretrained on one collection of recordings can generalize to new datasets, experimental paradigms, and even species.} Existing approaches are often designed for specific tasks and evaluated on a single dataset, making it unclear whether their learned representations are reusable for new calcium trace datasets. To tackle this gap, we present \textbf{CAPT}, a \textbf{C}ontinuous \textbf{A}utoregressive \textbf{P}opulation \textbf{T}ransformer for calcium population dynamics. CAPT models continuous calcium traces directly through a continuous patch tokenization strategy and is trained autoregressively, enabling end-to-end pretraining and adaptation to diverse downstream tasks. We first pretrain CAPT on a large-scale mouse calcium imaging dataset and evaluate its transferability across independent mouse, larval zebrafish, and \textit{C. elegans} datasets collected by different laboratories. In these transfer settings, the pretrained backbone is frozen and only adaptation modules are updated. Across neural population forecasting and behavior decoding tasks, CAPT consistently outperforms specialized and general-purpose baselines. Alongside predictive performance, multimodal analyses using NeuroPAL annotations in \textit{C. elegans} datasets show that CAPT embeddings form a shared functional space across datasets and capture anatomical cell-identity-related structure. These results suggest that the continuous autoregressive modeling opens up possibilities for a simple route towards general-purpose neural foundation models for calcium imaging, which can generalize across datasets, experimental paradigms, and species. Code is available at https://github.com/TSuXinH/CAPT.
Xinhong Xu, Yimeng Zhang, Yuanlong Zhang
School of Life Sciences Tsinghua University Beijing, China
Large language models (LLMs) achieve strong performance across many tasks but rely on dense multiply-accumulate (MAC) operations during inference, resulting in high energy cost. Spiking neural networks (SNNs) offer an event-driven alternative in which synaptic integration uses lightweight accumulation. However, spike-driven LLM inference remains difficult because outlier-heavy activations typically require long firing windows or auxiliary non-spiking paths. We propose QuantaSpike, a short-window spike-driven quantization framework for LLMs built around Logarithmic Ternary Integrate-and-Fire (LTIF) neurons. LTIF uses ternary events with power-of-two membrane-response quanta, improving the information represented by each firing step while retaining shift-ACC-compatible computation. QuantaSpike combines this neuron with group-adaptive gain and selective outlier admission: normal values use residual LTIF steps, whereas admitted outliers receive one additional onset spike before entering the same residual dynamics. Across OPT and Llama-2, QuantaSpike achieves state-of-the-art or competitive perplexity and zero-shot accuracy among spike-driven LLM quantization methods. It also transfers to newer dense LLMs, remaining close to the FP16 reference on Llama-3-8B and Qwen3-8B under the same four-step firing window. Analytical linear-energy projections show that QuantaSpike reduces the energy of one linear transformation by about 80.0% on OPT models and 67.1% on Llama-2 models relative to SpikeQuant, providing an accurate and energy-efficient spike-driven path for LLM inference.
Bang Hu, Guowei Zhu, Changze Lv +4
School of Computer Science, Fudan University, Shanghai, China
Robust and accurate neural decoders are integral to neurotechnologies such as brain-computer interfaces and closed-loop experiments. Recent work has shown that tokenizing neural data at the spike level facilitates multi-session pretraining and delivers state-of-the-art decoding performance. However, current spike-based models are restricted to supervised learning (SL), limiting training to datasets with paired behavioural labels. To address this limitation, we introduce MOJO (Masked autOencoder-based JOint training), a training framework for spike-tokenizing models that jointly leverages self-supervised learning (SSL) via masked autoencoding and SL objectives. We evaluate MOJO on three spiking datasets spanning monkey motor cortex during reaching tasks and multi-regional mouse recordings during vision and decision making tasks, demonstrating superior performance over purely SL-trained models. This improvement is especially pronounced when training with limited labelled data, particularly in few-shot finetuning, where only a small amount of labelled data from a new session is available. Incorporating SSL also yields more interpretable neuronal representations, improving performance on brain region classification and spike-statistics prediction without explicit optimization for these tasks. We further show that MOJO generalizes beyond spiking data to human electrocorticography during speech, where it continues to outperform purely SL-trained models and achieves performance comparable to neuro-foundation models (NFMs) designed specifically for continuous signals. Overall, augmenting spike-tokenizing models with SSL improves performance in label-impoverished settings and enables the use of unlabelled data across various tasks and species, while generalizing to other neural modalities. These results suggest a path towards more flexible and scalable data usage when training NFMs.
Ximeng Mao, Nanda H. Krishna, Avery Hee-Woon Ryoo +2
Mila – Quebec AI Institute · Université de Montréal · Canada CIFAR AI Chair