State Space Models
Also known as SSM
Momentum
28 papers in the last four weeks, up 250% on the four weeks before. 0.3% of all new papers.
Latest papers 181
Electromyography-to-Speech (ETS) synthesis is typically a non-linear, chaotic dynamical system. However, no prior work has studied the chaotic behavior of ETS synthesis to date. Yet, prior works strictly rely on standard reconstruction metrics with parameter-heavy transformers that systematically over-smooth natural acoustic dynamics. To close this gap, for the first time, we propose a chaos-inspired Lyapunov-derived activation function (LyRIC) with two novel chaotic loss functions, Lyapunov Exponent Regularization and Multi-Scale Detrended Fluctuation Analysis, to explicitly capture the deterministic chaos of human phonation. In addition, we introduce a compressed novel encoder, GLAME, which synergizes global Mamba state-space modeling with localized GLARE attention. We comprehensively perform frame-level acoustic evaluation in a multilingual and multi-speaker setup using English and Mandarin datasets. The proposed system outperforms the established baseline with a 4.69x increase in objective intelligibility (STOI: 0.61 vs. 0.13) and a 2.08x improvement in spectral reconstruction (LSD: 1.08 vs. 2.25). Importantly, this improvement is achieved with 73.49% fewer parameters (14.34M vs. 54.10M), establishing a new baseline for ETS synthesis. To the best of our knowledge, this is the first work demonstrating that integrating non-linear chaotic physics into neural networks yields superior yet compact inductive biases for real-time ETS synthesis.
MaRK: Markov-adapted Recurrent Kernels for Dynamic Operator Conditioning in State Space Models
State Space Models (SSMs) offer an efficient alternative to Transformers for sequence modeling, yet conditioning pre-trained SSMs for iterative generation typically operates outside the recurrent operator, through input injection or activation modulation. While such mechanisms expose the model to conditioning information, they leave the underlying temporal dynamics fixed. We introduce MaRK (Markov-adapted Recurrent Kernels), a dynamic operator-conditioning framework that maps context vectors directly into bounded modulations of a frozen SSM's recurrence (), read-in (), read-out (), skip (), and discretization () parameters. Viewed through the lens of LPV-SSM systems, MaRK induces a context-indexed family of Markov parameter sequences, allowing each diffusion timestep to reshape the model's input-output memory kernel. We instantiate MaRK on a frozen 111M-parameter Hydra SSM backbone and study three adapter geometries: Hypernet, Chebyshev polynomial, and Discrete Cosine Transform kernels. Since these adapters modify the Markov parameter sequence through low-rank auxiliary maps on the frozen backbone, parameter-efficient fine-tuning arises as a structural consequence of the adaptation mechanism itself, requiring only 6.3--11M trainable auxiliary parameters to transition from a bidirectional objective to an iterative diffusion regime. The bounded recurrence parameterization further yields an analytic Affine Quadratic Stability certificate for the modulated recurrence. Through synthetic LPV recovery experiments and Markov-operator diagnostics, we show that MaRK recovers coordinate-invariant temporal operators under matched assumptions and produces distinct, stable timestep-conditioned memory profiles. Empirically, the Chebyshev variant yields the strongest performance, achieving an average validation loss of 2.55, followed by the DCT (2.59) and Hypernet (3.77) geometries.
WaveGSSM: Graph Wave State Space Models for Propagating Spatio-Temporal Patterns
Spatio-temporal graph models typically encode each snapshot with a GNN and then connect the resulting representations through a temporal module. This space-then-time design is effective, yet it does not explicitly represent how a pattern moves across the graph. We show empirically that, for a propagating process, the same present field can lead to different futures when its recent rate of change differs, motivating an explicit representation of motion in the predictive state. We introduce WaveGSSM, a second-order graph state-space model that maintains two coupled latent states at each node, one for the current pattern and one for its temporal rate of change. A graph-wave transition updates the motion state through graph interactions and uses it to advance the pattern state, coupling spatial propagation and temporal evolution within a single rollout. We evaluate WaveGSSM on four temporal-graph benchmarks and global weather forecasting. It consistently achieves the best mean performance across the temporal-graph benchmarks and reduces the geopotential RMSE by 20.2% on average for 1- to 5-day weather forecasts relative to a backbone-matched snapshot model, while better preserving large-scale atmospheric patterns.
A Structured State Space Sequence Model for Multi-Class Classification of Malware
By 2030, Internet of Things (IoT) devices are projected to reach 40 billion, with fast-paced technological advancements in fields such as industry, healthcare, agriculture, automobiles, and building/home automation systems. This expansion has created a large attack surface for cybercrime, as the majority of these devices open the door for cybercriminals to exploit vulnerabilities, as they lack adequate built-in security. Cybercriminals launch malware attacks to compromise systems or steal sensitive data, and once a system is compromised, a ransom is typically demanded for its release. Current cybersecurity measures in place are being outpaced by the rapid growth of the IoT, which is accompanied by a subsequent growth in malware variants being created per day. Recognizing this pitfall, this research examines and proposes a novel approach to malware detection and classification to safeguard devices from further attacks and make IoT systems more robust and secure. The framework proposed utilizes a Structured State Space Sequence (S4) model, which discretizes sequences of malware samples in a sequence and captures long-range dependencies, essentially identifying the "cause" and "effect" hidden within malware execution flow. This study presents two novel contributions: the first empirical application of the S4 model for malware analysis, and a comprehensive comparison of its performance against other deep learning architectures, laying the stepping stone for future research in this new paradigm.
Action-On-Item Preference Flow: A Shared Event Schema for Predictive and Generative Personalization
A user's movie, news, and dialogue histories differ in their native actions and outputs, yet each interaction supplies evidence that can update user memory. We study whether these histories can train one reusable update mechanism. An action-on-item schema pairs a mapped interaction role with a content embedding, allowing shared update parameters to operate on separate user states. We establish invariance to native relabeling, bounded state changes under item-embedding perturbations, and a pooled-training bound under explicit compatibility conditions. The Multi-Timescale State Hypothesis (MTSH) specifies how this evidence enters, persists, and is consumed; PerTIDE implements it with action gating, three state-space traces, fusion, and command-conditioned readout. On PENS, the same history encoder supports both next-news prediction and personalized headline generation. In a controlled PENS-to-MovieLens experiment, a frozen source-trained core exceeds an identically structured random core by 15.23 MRR points after fitting the same target consumer. On MIND, PerTIDE retains a 4.12-point MRR advantage over a same-input three-branch state-space control. Action, readout, and trace interventions identify complementary contributions to these gains. Together, the theory and experiments support learning history updates across compatible sources and reusing them through predictive and generative consumers.
CAMOS: Coupled Oscillatory State-Space Model for Multimodal Clinical Time-Series
Longitudinal clinical cohorts are multimodal, irregularly sampled and pervasively incomplete: in ADNI, positron emission tomography and cerebrospinal fluid assays are absent from roughly half of all visits. Linear state-space models handle irregular sampling gracefully but treat a missing modality by masking the input, leaving the transition operator untouched. We prove that this is a representational limitation: the latent state of any linear state-space layer whose transition operator does not depend on the availability pattern is an additive function of the availability indicators, so no such layer can represent an interaction between two modalities being jointly present or jointly absent. We propose CAMOS, which gives each modality a bank of second-order oscillators coupled through a matrix that sits inside the differential equation and is gated by availability, so the transition operator itself becomes a function of which measurements were taken. Coupling invalidates the analysis of uncoupled oscillatory models, and we restore it: a per-channel Gershgorin budget makes the effective stiffness positive definite uniformly over all availability patterns and all gaps, an energy argument charges amplification to availability transitions rather than sequence length, and a channel factorization preserves exact associative parallel scans. On ADNI, CAMOS outperforms uncoupled oscillatory state-space models and clinical fusion models on same-visit staging, landmark prediction and longitudinal forecasting, and under zero-shot transfer to OASIS-3 it is the only model that avoids collapse to the majority class.
SCORE-LM: State-Space Radar Representations with Language Models for Fault Diagnosis
Radar hardware faults threaten automated perception, motivating accurate, compact diagnosis and understandable maintenance guidance. We introduce SCORE-LM, which couples a small scatterer-conditioned operator-response encoder (SCORE) to an adapted local language model. SCORE combines self-referenced complex trajectories, physical descriptors, and a selective state-space branch, with source-only self-supervision and directional fault inference. On eight capture-excluded Rad-R fault recordings, it achieves state-of-the-art performance within the evaluated nine-model comparison: 88.39% mean capture recall and 88.20% four-fault macro-F1 at ten frames. Its 39,520 radar inference coefficients are 119.7 times fewer than RadrNet-DS-CI's, while recall is 15.56 percentage points higher than this strongest competitor. In a separate low-label protocol, SCORE reaches 71.58% recall with one labeled source window per class. A nonlinear projector converts four frozen fault similarities into five soft tokens, linking compact diagnosis to class-conditioned maintenance guidance. On 75 development questions covering 24 radar windows, language adaptation raises correct-fault answers from 45 to 62 (60.0% to 82.7%) relative to removing the co-trained adapters, while retaining the same projector. SCORE-LM thus combines a compact radar specialist with a language interface for communicating fault-specific inspection guidance.
Riccati State Space Models: Non-iterative Parallelization for Nonlinear Sequence Modeling
State space models (SSMs) achieve efficient sequence processing because their affine state updates are closed under composition and can therefore be evaluated with an associative parallel scan. Nonlinear recurrent models can provide richer, state-dependent dynamics, but generally lose this compositional structure: parallel evaluation then requires iterative methods that repeatedly linearize and scan the recurrence. We ask, what state-dependent nonlinear dynamics can be designed to remain exactly composable? We answer by introducing RiccatiSSM, a nonlinear SSM, in which each state dimension follows an input-conditioned Riccati differential equation. Its quadratic state dependence makes the local Jacobian explicitly state-dependent, while its exact per-step flow under piecewise-constant inputs is a Möbius transformation. Since Möbius maps are closed under composition and compose through matrix multiplication, the complete nonlinear state trajectory can be evaluated exactly with a single associative parallel scan, without iterative linearization. We further derive a constrained parameterization that ensures bounded, contractive dynamics, and avoids poles in the fractional-linear state update. Across long-sequence classification, regression, and forecasting tasks, RiccatiSSM achieves competitive predictive performance while reducing runtime by compared to the nonlinear LrcSSM under matched architectures. These results demonstrate that state-dependent nonlinear dynamics can retain exact composability and be evaluated efficiently within a single parallel scan.
GT-PSSM: Unified Probabilistic Framework for Stochastic Dynamics Modeling and Dependency Learning in Multivariate Time Series Anomaly Detection
Multivariate time series anomaly detection (MTAD) is crucial for ensuring the safe and reliable operation of complex systems. Many existing methods learn normal patterns by training reconstruction or forecasting models on predominantly normal data. However, a large portion of these approaches rely on deterministic models and their associated point-wise output errors for anomaly scoring. Since real-world multivariate time series are inherently stochastic due to measurement noise and intrinsic system randomness, purely error-based scores can be unreliable, as large errors may arise from benign fluctuations rather than true anomalies. Probabilistic approaches address this limitation by quantifying uncertainty in model outputs. In particular, probabilistic state-space models (PSSMs) provide a principled framework by modeling stochastic system dynamics through latent state transitions and measurement noise via emission models. Despite this advantage, existing PSSM-based MTAD methods often struggle to capture long-range temporal dependencies and inter-variable dependencies, as they typically rely on noise-sensitive recurrent architectures and lack explicit cross-variable structure modeling. To address these limitations, we propose Graph-Transformer-Enhanced Probabilistic State-Space Model (GT-PSSM), a novel PSSM-based MTAD method that tightly integrates PSSM-based probabilistic modeling of stochastic dynamics with Graph Transformer-based learning of temporal and inter-variable dependencies. By jointly modeling stochasticity, long-range temporal dependence, and variable interactions within a unified probabilistic framework, GT-PSSM enables more robust anomaly detection.
ADPTNet: Adaptive with Prescriptive Timescales Non-Linear SSM for Sequence Modelling
A central aim of neuromorphic computing is to provide a viable alternative to highly energy-intensive Transformer-based AI. However, efficient alternatives struggle to capture the set of qualities that have secured the Transformer's status as the de facto standard in sequence modelling. Any realistic contender must be data-adaptive, able to capture long-range dependencies, and GPU-parallelisable, but also non-linearly recurrent to enable complex reasoning. Based on evidence suggesting the auditory cortex operates on fixed timescales, this work proposes the ADaptive with Prescriptive Timescales Network (ADPTNet) as a potential solution to achieving all four properties simultaneously. ADPTNet is built around local topological conjugates, obtained by a novel combination of linear attention and Riemannian optimisation, applied to static global dynamics. This enables non-linear yet predictable long-term behaviour. Dynamical systems theory proofs provide theoretical guarantees for the parametric control of ADPTNet's timescales (its Lyapunov spectrum). ADPTNet improves performance on Selective Copying over Hawk, the existing method balancing long-range memory and adaptability, while also improving state tracking over linear SSMs like Mamba. On sequential CIFAR-10, ADPTNet matches linear SSM accuracy and outperforms existing selective models (incl. the Transformer), using fewer parameters. We also introduce a neuromorphic SpikingADPTNet, which achieves a new state-of-the-art accuracy on the Spiking Speech Commands dataset (). Finally, ADPTNet's constant timescales enable two efficient, Jacobian-free extensions to the DEER parallel simulation algorithm (Conv and Forward DEER) that retain the same average convergence. Conv DEER adds no computational overhead beyond the network's forward pass and enables non-linear RNN parallelisation via iterated convolutions for the first time.
Correct then Forecast: Observer State-Space Models for Time Series Forecasting
Time series forecasting requires extrapolating the dynamics of an observed process beyond the last available measurement. Yet recurrent forecasting models typically treat observations as inputs that directly control their latent dynamics. It leads to a regime change when these observations become unavailable at prediction time. Following a state-estimation perspective, we introduce Observer State-Space Models (OSSMs), a class of recurrent models that interprets the observed input time series as measurements of an underlying autonomous dynamical system. OSSMs explicitly separate latent-state propagation from measurement assimilation: a single transition governs the dynamics across both context and forecasting intervals, while available observations correct the estimated state through an observer. This formulation naturally exposes classical control-theoretic properties, including observability and convergence of the state estimation error. We further show that conventional and recent SSMs can be recovered as particular instances of our OSSM framework, thereby providing a unified interpretation of their recurrent dynamics and revealing modeling inconsistencies. We perform experiments across several benchmarks showing that OSSM achieves substantial improvements while maintaining the same parameter count and training setup as the corresponding SSM baseline. These results support a simple principle for recurrent forecasting: observations should correct the estimated latent state, rather than control the dynamics used to propagate it.
Identifying the Predictable Drift of a Semimartingale from Marginal Laws
A special semimartingale admits a unique decomposition into a local martingale and a predictable finite-variation part . We consider the identification of when is observed only through repeated cross-sections. The estimand is then the projection of the sampled predictable compensator onto the observable feature filtration, namely the current state together with whatever randomness is shared across the population, so that at a fixed diffusion coefficient the marginal flow identifies the drift only up to a Markovian projection. If the drift is an affine functional of an observed lag window, the joint problem is a convex quadratic programme whose solution is the pseudo-panel regression of econometrics. Our principal concern is the case, which we believe not to have been treated before, in which the drift is the output of a hidden linear dynamical system whose dynamics are themselves to be identified from the marginals. The joint problem is then a bilinear quadratically constrained programme, which we solve to certified global optimality by spatial branch and bound; with unpenalised state disturbances and a drift basis growing with the grid it is NP-hard already in latent dimension one, by reduction from rank-one matrix approximation, whereas the complexity of the deterministic system at fixed latent dimension remains open. A block-coordinate decomposition offers a cheaper alternative. For the estimator itself, we obtain rates at a fixed mesh, separated into Monte-Carlo, estimation and grid contributions.
Baseline Shape Decides the Verdict: A Controlled Re-Examination of Ternary Language Models at 60K Parameters
Ternary (1.58-bit) weights are attractive for microcontroller-class language models, but the sub-1M-parameter regime rests mainly on isolated, single-seed comparisons. One prominent example reports that a routed ternary block (convolution, diagonal SSM and sparse attention mixed by a per-token router) beats a parameter-matched full-precision transformer by 22% at 60K parameters, attributing this to inductive bias. We re-run it under one fixed recipe, three seeds per cell, 98 byte-level runs on one laptop. (i) Baseline shape dominates: at a 16M-byte budget, param-matched transformers span 22.6% in validation loss purely by depth/width choice - far more than any architecture effect we measure there - and the best-shaped transformer ties the routed model, so the published margin is at least partly a baseline-shape effect; the ordering of shapes reverses with budget, so no single fixed shape can be trusted. (ii) At 130M bytes the routed model does win, by 22.2-24.0% over the three transformer shapes we evaluate there - but a plain gated diagonal-SSM block beats it by a further 9.1%, and the routed model's own router puts most of its weight on its recurrent pathway, so the gain does not require routing. (iii) The ternary penalty differs by architecture at the larger budget (+5.3% best transformer vs. +19.5% routed, +28.1% gated SSM), but we cannot attribute that to architecture alone: our transformers keep learned positional embeddings in full precision, 11-22% of their parameters, so they are less quantized than the models they are compared with. (iv) A 90/10 full-precision-then-ternary schedule beats all-ternary training, but only at a stage-2 learning rate about 10x the pretraining peak; at a conventional fine-tuning rate it looks 15.3% worse, reversing the conclusion. The from-scratch baseline was not itself learning-rate tuned, which bounds (iii) and (iv). Code and run logs released.
Physics and Data Driven Transformer-Mamba Framework for Flow Field
While deep learning accelerates expensive partial differential equation solving in computational fluid dynamics (CFD), existing methods like PINNs and FNOs often struggle with generalization, noise robustness, and physical consistency. We introduce the Transformer-Mamba for Flow Field (TM4FF) framework, a physics-constrained operator learning model with three key innovations: a Residual Wavelet Mamba (RWM) layer for feature denoising, a Transformer-based attention mechanism for enhanced feature fusion, and a physics-informed loss using Fourier derivatives to enforce the Navier-Stokes equations. Experiments on four CFD datasets show TM4FF achieves high accuracy and robust generalization across varying flow conditions.
Non-Commutative State Tracking with Input-Dependent Low-Rank Updates in Mamba-3
State tracking from sequential observations can require both retaining information and updating it by composing observed operations. We extend Mamba-3's diagonal transition with an input-dependent low-rank reflection term to support noncommutative state tracking, in which the order of operations matters. The rank-one update couples state coordinates along an input-dependent direction, enabling non-diagonal state transitions within a single Mamba-3 block. The extension preserves Mamba-3's exponential-trapezoidal discretization, rotary embeddings (RoPE), and readout. For training, we adapt chunkwise computation to parallelize the proposed recurrence within each chunk. Experiments cover group word problems with discrete inputs and a shell game with continuous observations, in which a policy is trained by behavioral cloning. Among the models selected for their strong performance under fixed timing, the proposed model maintains higher tracking success on longer swap sequences in the shell game with continuous observations and timing jitter. These experiments show that the proposed method achieves high accuracy on the evaluated non-commutative tracking tasks, improving on standard Mamba-3. The extension thus offers a Mamba-3-based approach to non-commutative state tracking.
When Labels Are Scarce: An Oscillatory State Space Model for Vibration Diagnosis
Machine fault diagnosis from vibration requires learning from scarce labelled fault recordings while meeting the computational constraints of edge devices for local inference. We introduce DualRes, a compact oscillatory state-space model that combines two complementary spectral views of vibration, capturing rapid changes and fine frequency structure. Time-aligned views are processed by selective oscillatory memory, which learns how long to retain temporal patterns. The encoder contains 39,528 parameters. We evaluate supervised learning across six bearing datasets and a gearbox benchmark, with an additional gearbox pilot. Recording-level splits and explicit accounting of labelled duration distinguish data efficiency from repeated exposure to correlated samples. On the main gearbox benchmark, DualRes achieves state-of-the-art performance among the nine evaluated methods at six of seven label budgets. With about six labelled seconds per class, it improves macro-F1 by 16.1 percentage points over the next strongest comparator. On the same benchmark, DualRes achieves a 1.44-fold recording-level speedup and a 24.8-fold reduction in checkpoint storage relative to a selective state-space baseline under matched hardware and runtime conditions. Bearing results reveal task-dependent trade-offs. These findings support oscillatory memory as a compact approach to vibration diagnosis under limited labelled exposure.
Full-Covariance Smoothing of Bayesian Neural Networks for Online Adaptation
A neural network's layers can be treated as time steps of a state-space model, turning Bayesian training into a smoothing problem: a forward pass propagates Gaussian moments through the network, and a backward Rauch--Tung--Striebel pass updates the weight posteriors in closed form. Such methods learn from each observation in a single pass, in an uncertainty-aware manner, and without gradient-based iterations or replay, which makes them well suited for online adaptation and data-efficient learning. Existing smoothing-based methods, however, are restricted to diagonal covariances across activations, discarding correlations between neurons. We overcome this limitation via a cross-covariance identity that enables full-covariance propagation through a network's nonlinear activations. We derive a one-step-per-layer smoother that approximates as Gaussian only each layer's affine output, and that applies both to deterministic systems with noisy observations and to stochastic systems described by output statistics. We demonstrate this method in non-stationary classification, online dynamics learning, and policy adaptation of a vision-language-action model, and find that it is generally more accurate than other smoothing-based methods.
Temporally Ordered Region-Token Mamba with Logit-Space Diffusion for Remote Sensing Change Detection
Remote sensing change detection requires both global reasoning across bitemporal images and precise localization of changed regions. However, dense attention is computationally expensive for high-resolution imagery, while conventional feature fusion and coarse decoding may inadequately separate genuine changes from appearance variations or preserve object boundaries. We present Bitemporal Mamba-Diffusion for Change Detection (BMD-CD), which combines temporally structured state-space modeling with logit-space diffusion refinement. BMD-CD converts deep bitemporal features into region tokens and arranges them in explicit temporal partitions before bidirectional state-space propagation. Its Bitemporal Ordered Mamba Operator enables long-range cross-temporal interaction with linear sequence complexity, while Orthogonal Feature Disentanglement forms a change-oriented output and a complementary rotated output using learned pairwise rotations and unchanged-region consistency. Multiscale decoding then produces coarse change logits, which are refined through a five-step Conditional Diffusion Decoder operating directly in logit space. Experiments on LEVIR-CD, WHU-CD, DSIFN-CD, CDD, and S2Looking demonstrate strong performance across diverse change-detection settings. BMD-CD achieves F1 scores of 93.7%, 96.0%, 97.8%, and 99.0% on the four standard benchmarks and improves 3-pixel Boundary-F1 to 87.7% and 91.4% on LEVIR-CD and WHU-CD, respectively. The full model requires 32.09 GFLOPs and 47 ms per 256 x 256 image pair, while also showing zero-shot transfer to ValaisCD and B-FLAIR-test. Our code is available at https://github.com/Aparup2139/Public_WACV/
PR-Smoother: Simulator-Preserving Non-Gaussian Smoothing for Data Assimilation
Many physical data assimilation (DA) workflows require smoothing methods that represent non-Gaussian posteriors over physical state variables, scale to high-dimensional simulators, train from observation windows alone, and remain compatible with calibration of the prescribed simulator. We introduce PR-Smoother, a simulator-preserving amortized smoother designed for this prescribed-simulator DA regime. Its key design principle is to keep the prescribed simulator explicit in both the evidence lower bound and the variational family: rather than learning replacement dynamics or a learned trajectory prior, PR-Smoother learns only future-conditioned corrections around the prescribed rollout. This yields an explicit non-Gaussian smoothing distribution over physical trajectories and supports joint state, parameter, and sensor-bias learning from observations alone. The variational family contains the exact smoother in deterministic and linear-Gaussian limits. Empirically, PR-Smoother captures multimodal posteriors in 4-dimensional Lorenz-96, remains accurate under ambiguous nonlinear observations and process noise in 40-dimensional Lorenz-96, and scales to joint state-parameter-bias inference in 16,384-dimensional Kolmogorov flow.
Comparing Latent Concept Formation in State Space Models and Transformers via Sparse Autoencoders
The quadratic scaling of Transformer self-attention has driven the adoption of sub-quadratic Selective State Space Models (SSMs) like Mamba, which compress past context into a fixed-size recurrent hidden state. This strict informational bottleneck raises a foundational question for mechanistic interpretability: do SSMs and Transformers learn fundamentally distinct latent representations? In this work, we employ Sparse Autoencoders (SAEs) to conduct a large-scale, feature-level correspondence analysis between Mamba-130m and Pythia-70m over a 10-million token corpus. Contrary to hypotheses predicting widespread architectural divergence, we find no evidence of systematic representational divergence between architectures: across the observed Jaccard distribution, 99.98% of Mamba features cluster toward the upper alignment boundary, providing preliminary feature-level support for the Universality Hypothesis. We further identify and qualitatively characterize this microscopic fraction (0.02%) of diverging features, finding patterns consistent with the hypothesis that the recurrent bottleneck selectively limits the parsing of rigid syntax rather than broad semantic ontology. We demonstrate that while Pythia's unconstrained attention permits the monosemantic decomposition of distinct formatting edge-cases, Mamba is forced to compress unrelated syntactical anomalies into polysemantic "junk drawer" neurons to preserve state capacity. Collectively, these results suggest that architectural routing mechanisms may have negligible impact on core semantic understanding, with representational divergence confined to extreme structural margins.
Pharmacokinetic State Space Models for Unbiased Prediction of Haemodynamic Collapse
An Intraoperative Hypotension (IOH) event is a frequent complication during administration of general anaesthesia with serious downstream consequences, yet clinical management remains reactive and not predictive. Existing predictive models, however, ignore drug infusion history as a valuable signal for prediction despite its direct pharmacological relevance. Our model achieves an Area Under the Receiver Operating Characteristic curve (AUROC) of 0.7360 and an Area Under the Precision-Recall Curve (AUPRC) of 0.1794, representing a 2.73-fold lift over the random guessing AUPRC baseline (0.0657), with the removal of propofol and remifentanil effect-site concentrations resulting in a 13.9% AUPRC drop compared to the full model. This is consistent with the hypothesis that pharmacokinetic trajectories encode impending haemodynamic changes before they manifest in the Mean Arterial Pressure (MAP). Additionally, this paper shows that training without lead-gap filtering degraded AUROC by 16.7%, empirically confirming that unfiltered models learn to detect ongoing hypotension rather than predict future events. Finally, a Mamba-based architecture achieves the aforementioned high prediction performance while maintaining a constant memory footprint across a range of sequence lengths, unlike the quadratic VRAM overhead typical of vanilla Transformers, making it the more practical choice for continuous intraoperative deployment.
REACT: A Fully Spiking State-Space Model for Real-Time Event-Driven Temporal Perception
Robotic systems operating in dynamic environments require visual perception that evolves continuously with the incoming sensory stream. Event cameras provide microsecond temporal resolution and asynchronous sensing, but most learning-based methods accumulate events into frames or temporal bins, introducing an integration delay that can limit fast reaction. Here we propose REACT, a fully spiking state-space model for event-driven temporal perception that processes raw events one by one, without temporal accumulation. REACT uses a complex-valued spiking neuron, C-SiLIF, whose continuous-time dynamics are driven by the physical inter-event interval, allowing its internal state to evolve at the temporal resolution of individual events. We evaluate REACT on gesture recognition and time-to-collision (TTC) estimation from full-field event streams, without a target bounding box or localization input. On EvTTC, REACT achieves a 9.59% relative TTC error with 4.6 ms end-to-end inference latency, within 0.15 percentage points of the best learned method while requiring no target prior. At the dataset's mean approach speed, this latency corresponds to only 4 cm of vehicle motion, compared with 1 m for the fastest competing learned method. REACT further supports anytime TTC prediction, zero-shot transfer to a different driving sequence, and INT8 quantization, reducing the estimated energy consumption from 18.5 to 2.8 mJ per 32,768 events. These results show that event-driven spiking state-space dynamics can provide low-latency, continuously updated temporal perception for reactive robotic systems.
Long-Context Demonstration Selection Using State Space Models
We study the problem of demonstration selection, which involves selecting a subset of examples for prepending to a query to a language model. This problem is closely related to in-context learning and language model inference. Since the inference cost of a transformer model scales quadratically with sequence length, the selection problem becomes especially challenging in a long-context scenario. In this paper, we tackle this problem by building on state space models (SSMs), which require only linear inference time given the input. Our approach involves two algorithms. The first learns a small set of SSMs through distillation of a (trained) transformer model. We partition all the layers into consecutive groups. Then for each group, we estimate a separate state space model to replicate the input-output behavior within the adjacent layers. Second, we map the distilled model outputs to a small set of tokens, and apply these embeddings for demonstration selection in downstream applications. We perform extensive experiments in both synthetic and real-world datasets to validate our approach. We demonstrate that the distilled SSMs only incur an approximation error of less than relative to the true output. In downstream evaluation, we show that on several text classification and reasoning tasks, our approach reduces FLOPs by and improves accuracy by relative to baseline demonstration selection methods.
On the Importance of Gating: Memorization vs. In-Context Learning in State Space Models
State Space Models (SSMs) have emerged as a compelling alternative to Transformers, enabling sequence modeling with constant memory and linear compute. Although SSMs exhibit reasonable performance and favorable computational characteristics, they continue to lag behind Transformers on tasks that require in-context learning and precise retrieval, slowing their adoption for large-scale language modeling. In this work, we demonstrate that both the success and failure of SSMs in these domains can be explained by studying the role of the gating mechanism, a prevalent component in modern recurrent networks. Specifically, we show through theory and experiments that this gating mechanism causes SSMs to first learn an in-weights "memorization" solution, while delaying, or even preventing, convergence to a correct in-context learning solution. Importantly, this happens even in cases where there are no fundamental limitations due to the architecture or its memory capacity. On the other hand, we find that gating is often beneficial for improving generalization to long sequence lengths. Our results illuminate the crucial role of the gating mechanism in shaping both the training dynamics and generalization of SSMs, and provide a basis for understanding and improving linear-time models.
Anatomy of Associative Recall in Fixed-State Recurrences: A Matched-State Decomposition, an Interference Wall, and a Curriculum That Breaks It
Fixed-state recurrences--linear attention and state-space models--are reported to lag behind attention on associative recall, but whole-architecture comparisons cannot say which ingredient is responsible. We decompose masked multi-query recall at a fixed state budget along three single-knob axes: a short causal convolution, the transition structure (rank-1 delta rule vs. diagonal), and decay. The convolution dominates (~+0.5 recall in both families under matched training): comparisons that pit convolution-free cells against a convolution-equipped Mamba measure the missing convolution, not the recurrence. The rank-1 transition beats its diagonal ablation by +0.19/+0.32 at 16/32 pairs, but the margin shrinks to +0.03 once both cells carry the convolution, and a state-matched Mamba-2 ties the unarmed rank-1 cell: no class claim survives. Cells that solve 32-pair recall degrade gracefully with load yet fall to chance retrieving 4 pairs from a distractor haystack--flat across lengths and transitions. Interference under sparse supervision, not capacity: a distance curriculum takes the unchanged architecture from 0.021 to 1.000. Training is a lock-in lottery--a seed either locks in or does not--and the curriculum is the lever. Lock-in rises from 1/10 to 7/10 (p=0.02); dense supervision adds nothing; at L=256 a shaped ramp reopens a boundary the uniform curriculum cannot (4/5 vs. 0/9); and at L=512, where the ramp collapses (0/6), gating it on measured accuracy locks in 6/6 (p=0.001). Bidirectional denoiser cells, reading the query before the haystack, show no measurable advantage over causal training (ten seeds), and collision-key retrieval needs two layers. Arming for recall is free on an S_5 state-tracking guardrail--the armed cell is significantly better at every depth (p<=0.0044). These replace "recurrent models are bad at recall" with a measured decomposition and two cheap interventions.
Fixed State, Long Reach: What a Constant-Size Cache Buys Block Diffusion at Scale
Diffusion language models decode tokens in parallel, but their bidirectional denoiser rules out the naive key--value (KV) cache behind fast autoregressive inference. Block diffusion restores caching by decoding block-by-block, and the block caches deployed on it so far are tied to attention: O(L)in memory and, if used as training-free retrofits, only an approximation of the model's computation. Both constraints can be overcome: sequence mixers that summarize finalized blocks into a reusable state support block caching, and the corresponding block-causal training objective makes the cache exact. We study this recipe at scale, pretraining three 3B block-diffusion denoisers (attention, mamba, and hybrid) on 300B tokens under one single-frontier objective and decoding all three through a single cached interface. Only the state-space cache is O(1) in sequence length: its memory and per-step latency stay constant at any context length, while an attention cache remains O(L). At 256k tokens (where attention has grown to 82GB and 29 ms/step), the Mamba cache delivers 4.3x lower latency, 11x less memory, and 2.6x higher single-stream throughput; and because that footprint is constant it scales with batch as well, reaching 14x the aggregate throughput, where attention cannot run beyond a single stream. The same linear-state bias lets the Mamba and hybrid backbones keep retrieving out to 8-16x their training length, whereas attention's retrieval collapses at 2x, at no measured quality cost.
OphBiWSSD: Scaling Temporal Action Localization in Ophthalmic Surgeries with Bidirectional Weight-tied State Space Duality
High-frequency surgical maneuvers in ophthalmology necessitate high-fidelity temporal modeling, yet characterizing long-range procedural dependencies remains computationally prohibitive for attention-based architectures. Existing models often require aggressive temporal downsampling, which compromises the detection of fine-grained action boundaries and instrument-tissue interactions. To address these scalability constraints, we present OphBiWSSD, a framework that reformulates surgical temporal action localization leveraging Bidirectional State Space Duality. By employing a weight-tied selective scan mechanism that incorporates both preceding and succeeding surgical contexts, our approach facilitates the global synthesis of non-causal temporal cues with linear complexity. This streamlined architecture is well-suited to capture the bidirectional dependencies present in ophthalmic workflows, effectively bridging the gap between local boundary precision and long-range procedural context without incurring the quadratic memory overhead of traditional Transformers. Extensive experiments on the OphNet benchmark demonstrate that OphBiWSSD achieves state-of-the-art temporal localization performance, with mean Average Precisions of 44.42% on phases and 43.08% on operations, surpassing the baselines by 6.80% and 6.66%, respectively. Empirical validation indicates that our approach ensures precise temporal localization and offers a computationally viable pathway for deploying surgical intelligence systems in clinical environments. The code is publicly available at https://github.com/yo3nglau/OphBiWSSD.
PAC-Bayesian Bounds for Learning Partially Observed Stochastic Linear Time-Invariant State-Space Systems with Inputs and Sub-Gaussian Noise
In this paper we derive a Probably Approximately Correct (PAC)-Bayesian error bound for partially observed linear time-invariant (LTI) stochastic dynamical systems in state-space form with inputs and sub-Gaussian noise. Such bounds are widespread in machine learning, and they are useful for characterizing the predictive power of models learned from finitely many data points. The bound derived in this paper relates the expectation of prediction errors with the prediction error generated by the model on the data used for learning. In addition, we show that it can also be used to derive bounds for the parameter estimation error. In turn, this allows us to provide finite-sample error bounds for the prediction error and parameter estimation error for a wide class of system identification algorithms. Furthermore, as LTI systems are a sub-class of recurrent neural networks (RNNs), these error bounds could be a first step towards PAC-Bayesian bounds for RNNs.
Global Divergence, Local Convergence: Representation Geometry in SSMs and Transformers
Recent state-space models (SSMs) such as Mamba achieve language modeling performance comparable to transformers despite relying on fundamentally different architectures. This raises an important question: how do these structural differences influence the geometry and functional nature of their internal representations? We study this question through a multi-scale analysis of representations in transformers, SSMs, and hybrid architecture. First, we find that SSMs distribute their representational information evenly across all dimensions, whereas transformer representations are heavily dominated by a single principal direction. By evaluating hybrid architectures, we observe that the representation space becomes increasingly skewed toward a single dominant direction after each attention layer. Next, we explore how the different geometric spread of representations impacts representational capacity through compressibility. Surprisingly, we find that despite their contrasting geometric structures, both architectures exhibit tightly matched effective capacities. We further investigate whether this skewed geometry affects how concepts are encoded. Using rank-constrained probes, we demonstrate that both architectures encode concepts in subspaces of surprisingly similar dimensionality. Furthermore, we demonstrate that the transformers' dominant principal direction does not inherently encode more conceptual information. Finally, we zoom in and examine the alignment between manifolds, either by analyzing representations of specific topics or by looking at the nearest neighborhoods of tokens, and find that they are highly aligned. Ultimately, our analysis suggests that while transformers and SSMs induce different usage of latent space, they display a striking functional convergence at the level of local semantic manifolds.
On the Recall Scaling Laws in Mamba: A Theoretical and Mechanistic Study via Hashing
Associative Recall (AR) is the cognitive ability to learn and retrieve links between items in memory. In NLP, AR is used as a benchmark for evaluating the in-context memory capacity of architectures such as Mamba, and has been found to strongly correlate with language modeling performance. This paper explores AR from the perspective of mechanistic interpretability, aiming to reverse-engineer the exact internal algorithm used by Mamba to perform recall. Our key insight is that Mamba performs recall by implicitly learning linear hash functions, and we identify the low-level circuit that enables this behavior. Building on these findings and inspired by theoretical tools in similarity-preserving hashing, such as the Johnson-Lindenstrauss lemma, we develop a theoretical framework for analyzing AR, which we term Recall Scaling Laws. Given the vocabulary size and the number of facts in context, this framework allows us to (1) predict the embedding and state dimensions required for Mamba to achieve perfect recall, (2) predict recall success probability given the model dimensions, and (3) analyze multi-layer models and multi-head SSM patterns. Empirical results show that our theoretical findings are accurate and predictive, offering insights into how AR capacity scales with vocabulary, state, embedding size, and architecture.