math.OCMar 13, 2026

State-space models through the lens of ensemble control

Authors: Ye Feng, Jianfeng Lu

Abstract

State-space models (SSMs) are effective architectures for sequential modeling, but a rigorous theoretical understanding of their training dynamics is still lacking. We formulate continuous-time SSM parameter-path optimization as an ensemble optimal control problem: each input sequence generates a corresponding state trajectory in the ensemble, while the trainable parameter path forms a common control shared by all trajectories. Within this formulation, model evaluation is represented by the forward state equation and backward sensitivity propagation by the corresponding adjoint equation. We show that the Hamiltonian gradient with respect to the control variable represents the negative first-variation density of the reduced ensemble objective. Using this representation, we analyze a Bregman mirror-descent scheme for the continuous-time ensemble objective; in Euclidean geometry this reduces to functional projected gradient descent. State-adjoint and Hamiltonian-gradient stability yield two-sided relative curvature of the reduced objective and relative convexity under sufficient regularization. We further show that in the relatively convex regime the optimal-control problem admits a minimizer, which is unique under relative strong convexity. For SSMs, the explicit stability-generated defect κstabSSMκ_{\mathrm{stab}}^{\mathrm{SSM}} gives an O(1/k)O(1/k) objective rate at τ=κstabSSMτ=κ_{\mathrm{stab}}^{\mathrm{SSM}} and geometric convergence to the unique minimizer for τ>κstabSSMτ>κ_{\mathrm{stab}}^{\mathrm{SSM}}, under the stated stepsize condition.

Explore similar work

May 16, 2025cs.LG

Regularity and Stability Properties of Selective SSMs with Discontinuous Gating

Selective State-Space Models (SSMs) such as Mamba have become central to long-sequence modeling. Still, their stability is poorly understood: their state-space coefficients are modulated online by a token-dependent gating signal, making the recurrence neither linear time-invariant nor classically nonlinear. We study continuous-time selective SSMs through passivity, dissipativity, and Input-to-State Stability (ISS), explicitly separating the selection signal x(⋅)x(\cdot) from the driving input u(⋅)u(\cdot). We obtain four results: exponential forgetting under strict dissipativity; a canonical AUCloc\mathrm{AUC}_{\mathrm{loc}} quadratic storage for the frozen-selection subsystem that accommodates discontinuous gating; a parametric LMI together with universal kernel constraints and "irreversible forgetting" under universal quadratic storage; and sufficient conditions for global ISS uniformly over admissible selection schedules. We then bridge to practice by deriving a sampled block LMI for the Mamba selective-scan core, which is used as a differentiable training-time regularizer. Across seven standard time-series datasets and four prediction horizons, the regularizer reduces sampled Mamba-core LMI violations by roughly 92%92\% in 28/2828/28 pairs at a clean-MSE cost of less than 0.018%0.018\%. It improves internal Mamba passivity and state-norm diagnostics under injected perturbations. Our results turn classical control-theoretic tools into verifiable structural and training criteria for selective SSMs, while honestly scoping which guarantees transfer to a deep selective-scan architecture.
Jun 29, 2026cs.LG

MuonSSM: Orthogonalizing State Space Models for Sequence Modeling

State space models (SSMs) have emerged as efficient linear-time alternatives to attention for long-sequence modeling. However, existing SSMs often suffer from instability and memory degradation over extended horizons due to poorly conditioned first-order updates and unbalanced update geometry. We introduce MuonSSM, a general framework that stabilizes SSM training by explicitly conditioning the geometry of memory updates rather than the recurrent transition matrix. MuonSSM augments SSMs with a momentum-based pathway and a lightweight Newton Schulz transformation on low-rank input injections, yielding bounded and spectrally conditioned updates while preserving parallel scan complexity. Theory shows that MuonSSM improves gradient propagation, mitigates spectral amplification, and enriches memory representations over long horizons. Extensive experiments across language, vision, and time-series benchmarks show consistent gains in accuracy, robustness, and long-context performance when integrated into diverse SSM backbones. These results establish geometric conditioning of updates as a principled pathway to stable, scalable sequence modeling.
Dec 22, 2025cs.LG

Lag Operator SSMs: A Geometric Framework for Structured State Space Modeling

Structured State Space Models (SSMs), which are at the heart of the recently popular Mamba architecture, are powerful tools for sequence modeling. However, their theoretical foundation relies on a complex, multistage process of continuous-time modeling and subsequent discretization, which can obscure intuition. We introduce a direct, first-principles framework for constructing discrete-time SSMs that is both flexible and modular. Our approach is based on a novel lag operator, which geometrically derives the discrete-time recurrence by measuring how the system's basis functions undergo what we call a domain expansion from one timestep to the next. The resulting state matrices are computed via a single inner product involving this operator, enabling a modular design space for creating novel SSMs by flexibly combining different basis functions and time-warping schemes. To validate our framework, we demonstrate that a specific instance exactly recovers the recurrence of the influential HiPPO model. Numerical simulations confirm our derivation, providing new theoretical tools for designing flexible and robust sequence models.