cs.LGMar 14, 2026

PDE-SSM: A Spectral State Space Approach to Spatial Mixing in Diffusion Transformers

Authors: Eshed Gal, Moshe Eliasof, Eldad Haber

Organizations: University of British Columbia, Vancouver, BC Canada · University of Cambridge, Cambridge, United Kingdom

Abstract

The success of vision transformers-especially for generative modeling-is limited by the quadratic cost and weak spatial inductive bias of self-attention. We propose PDE-SSM, a spatial state-space block that replaces attention with a learnable convection-diffusion-reaction partial differential equation. This operator encodes a strong spatial prior by modeling information flow via physically grounded dynamics rather than all-to-all token interactions. Solving the PDE in the Fourier domain yields global coupling with near-linear complexity of O(Nlog⁡N)O(N \log N), delivering a principled and scalable alternative to attention. We integrate PDE-SSM into a flow-matching generative model to obtain the PDE-based Diffusion Transformer PDE-SSM-DiT. Empirically, PDE-SSM-DiT matches or exceeds the performance of state-of-the-art Diffusion Transformers while substantially reducing compute. Our results show that, analogous to 1D settings where SSMs supplant attention, multi-dimensional PDE operators provide an efficient, inductive-bias-rich foundation for next-generation vision models.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Pixel-Space Diffusion Transformers

    Jul 20, 2026Renye Yan, Jikang Cheng, You Wu +8Diffusion TransformersVisual Tokenizers

  2. Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer

    Nov 24, 2025Haoyu Wu, Jingyi Xu, Qiaomu Miao +2Diffusion TransformersRotary Position Embedding