cs.CVSep 29, 2026

Parameterized Stripe Attention for Efficient Video Generation

Authors: Xingyu Jia, Baole Ai, Ang Wang, Kang Zhao, Yong Li

Organizations: Alibaba Group

Abstract

Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the computationally expensive full spatio-temporal attention. While sparse attention methods offer potential solutions, existing approaches face an inherent flexibility--efficiency dilemma: predefined masks lack the flexibility to capture diverse attention patterns, while runtime-determined masks introduce overheads and sacrifice hardware efficiency. We identify the lack of a unified structural characterization of DiT attention as a key limitation of existing methods, and establish that video DiT attention exhibits \textbf{periodic diagonal stripe structures} along both temporal and spatial dimensions. To formally encode these structured patterns within a single efficient kernel, we present {\bf PSA}, a parameterized stripe attention that formalizes the observed stripe regularity, unifying diverse attention patterns for efficient mask generation. This unified representation enables a single hardware-efficient CUDA kernel to process all sparse patterns, achieving FlashAttention-3-level Model FLOPs Utilization. To determine optimal sparsity configurations, we propose a training-free offline search algorithm that automatically maximizes sparsity under a specified error tolerance for each attention head. Experiments on HunyuanVideo and Wan~2.1 demonstrate that PSA achieves 1.57×\times and 1.37×\times end-to-end speedups over FlashAttention-3 baselines, with acceptable visual quality degradation.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ScalingAttention: Discovering Intrinsic Sparse Attention Topology for Video Diffusion Transformers

    Jun 22, 2026Ruiliang Zhou, Xuecheng Wu, Kang He +6Dynamic Sparse AttentionDiffusion Transformers

  2. Veda: Scalable Video Diffusion via Distilled Sparse Attention

    May 28, 2026Shihao Han, Hao Yang, Xinting Hu +3Video Diffusion ModelsDynamic Sparse Attention