cs.LGOct 6, 2026

MaRK: Markov-adapted Recurrent Kernels for Dynamic Operator Conditioning in State Space Models

Authors: Syed Ibrahim Omer, Ginny Y. Wong. Xiangyu Zhao

Organizations: Department of Data Science City University of Hong Kong · NVIDIA AI Technology Center NVIDIA, Santa Clara, USA

Abstract

State Space Models (SSMs) offer an efficient alternative to Transformers for sequence modeling, yet conditioning pre-trained SSMs for iterative generation typically operates outside the recurrent operator, through input injection or activation modulation. While such mechanisms expose the model to conditioning information, they leave the underlying temporal dynamics fixed. We introduce MaRK (Markov-adapted Recurrent Kernels), a dynamic operator-conditioning framework that maps context vectors directly into bounded modulations of a frozen SSM's recurrence (AA), read-in (BB), read-out (CC), skip (DD), and discretization (ΔΔ) parameters. Viewed through the lens of LPV-SSM systems, MaRK induces a context-indexed family of Markov parameter sequences, allowing each diffusion timestep to reshape the model's input-output memory kernel. We instantiate MaRK on a frozen 111M-parameter Hydra SSM backbone and study three adapter geometries: Hypernet, Chebyshev polynomial, and Discrete Cosine Transform kernels. Since these adapters modify the Markov parameter sequence through low-rank auxiliary maps on the frozen backbone, parameter-efficient fine-tuning arises as a structural consequence of the adaptation mechanism itself, requiring only 6.3--11M trainable auxiliary parameters to transition from a bidirectional objective to an iterative diffusion regime. The bounded recurrence parameterization further yields an analytic Affine Quadratic Stability certificate for the modulated recurrence. Through synthetic LPV recovery experiments and Markov-operator diagnostics, we show that MaRK recovers coordinate-invariant temporal operators under matched assumptions and produces distinct, stable timestep-conditioned memory profiles. Empirically, the Chebyshev variant yields the strongest performance, achieving an average validation loss of 2.55, followed by the DCT (2.59) and Hypernet (3.77) geometries.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MuonSSM: Orthogonalizing State Space Models for Sequence Modeling

    Jun 29, 2026Thai-Khanh Nguyen, Ngoc-Bich-Uyen Vo, Thieu N. Vo +2State Space ModelsSequence Modeling

  2. Lag Operator SSMs: A Geometric Framework for Structured State Space Modeling

    Dec 22, 2025Sutashu Tomonaga, Kenji Doya, Noboru MurataState Space ModelsState Space