cs.CVSep 30, 2026

PARK: Accurate Block Retrieval for Sparse Attention in Video Diffusion Transformers

Authors: Yun Dai, Jiarui Wen, Huiping Zhuang, Cen Chen, Ziqian Zeng

Organizations: South China University of Technology

Abstract

Diffusion Transformers (DiTs) have become a dominant architecture for video generation, but their efficiency is limited by the quadratic complexity of full attention. Sparse attention reduces this cost by retrieving important blocks and computing attention only within them, but inaccurate retrieval can either degrade generation quality or yield unnecessary computation. We identify two retrieval mismatches in methods that retrieve blocks using the averaged representations of query and key blocks: (i) query-side aggregation mismatch, where averaging queries before Softmax fails to preserve their individual attention preferences, and (ii) key-side clustering metric mismatch, where standard Euclidean clustering in the original key space can group keys with dissimilar QK scores under the current query, so their average representation may not accurately represent how the current query scores individual keys. These mismatches can lead to inaccurate block retrieval. To address these mismatches, we propose PARK, a training-free sparse attention method for accurate block retrieval. PARK retains every original query, independently normalizes its attention over key blocks, and then averages these distributions within each query block. It also uses information from the current queries to transform keys before clustering, so that keys receiving similar QK scores are grouped together. A fused GPU kernel further reduces the overhead of block retrieval. Experiments on HunyuanVideo and Wan demonstrate that PARK improves block retrieval accuracy and preserves generation quality while accelerating inference, achieving the best quality-efficiency trade-off among the compared sparse attention methods.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video Generation

    Apr 20, 2026Haoyue Tan, Shengnan Wang, Yulin Qiao +5Clustering

  2. SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference

    Aug 4, 2026Shanghao Liu, Renze Chen, Size Zheng +4Dynamic Sparse AttentionSparsity