cs.LGApr 16, 2026

On the Expressive Power and Limitations of Multi-Layer SSMs

Authors: Nikola ZubićQian LiYuyi WangDavide Scaramuzza

Organizations: Robotics and Perception Group University of Zurich · Shenzhen International Center For Industrial And Applied Mathematics Shenzhen Research Institute of Big Data · Tengen Intelligence Institute2026 CRRC Zhuzhou Institute

Abstract

We study how depth, finite precision, state dimension, and chain-of-thought (CoT) affect the expressive power of multi-layer state-space models (SSMs). For the explicit-table KK-function-composition problem, a canonical benchmark for sequential information propagation, we prove that any LL-layer SSM solving (L+3)(L+3)-function composition must satisfy d2p=Ω(N/L3)d^2p=Ω(N/L^3), where dd is the state dimension and pp is the per-scalar precision. Conversely, KK-function composition is solved exactly by a (K+1)(K+1)-layer generalized SSM with d=1d=1 and p=Θ(logN)p=Θ(\log N). This gives a worst-case depth hierarchy for this formal problem family. We then distinguish post-input reasoning, in which all thought tokens are generated after the input, from input-interleaved reasoning, in which thought tokens may be inserted while the input stream is being read. Post-input reasoning does not circumvent our communication-based lower-bound pipeline, whereas input-interleaved reasoning admits bidirectional simulations with general deterministic one-pass streaming algorithms at the granularity of persistent memory. Finally, width and precision are not interchangeable under exact step-preserving simulation in the base affine-state model, but become interchangeable through the streaming-memory characterization once input-interleaved reasoning is allowed.

Explore similar work

CardsList
  1. Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models

    May 18, 2026Aleksandar Terzić, Francesco Carzaniga, Nicolas Menet +4State Space ModelsFlash