cs.CVSep 28, 2026

Persistence Forcing: Exploiting Feature Specialization in Pixel-Space Diffusion

Authors: Chong Wang, Zixuan Fu, Shiqi Huang, Siyuan Yang, Hao Cheng, Bihan Wen

Organizations: Nanyang Technological University · KTH Royal Institute of Technology · Hebei University of Technology

Abstract

Pixel-space diffusion Transformers (DiTs) directly operate on high-dimensional visual data, yet their hidden representations typically undergo uniform refinement across depth. Natural images, however, are inherently organized at different levels of granularity. Global structure can often be represented compactly, whereas local textures and fine details require richer representations. Motivated by this, we introduce heterogeneous refinement in pixel-space DiTs, assigning different feature groups distinct refinement budgets across depth. Consequently, an ordered feature specialization emerges: sparsely refined features predominantly encode global visual structure, whereas more frequently refined features increasingly specialize toward localized, high-frequency details. We refer to these two groups as persistent and active features, respectively. Building on this emergent specialization, we introduce Persistence Forcing (PerF), which explicitly exploits this persistent--active feature organization for pixel-space image generation. This enables persistent features to continuously condition actively refined features, allowing stable global information to guide the ongoing refinement of finer visual details. During generative sampling, this interaction further induces a meaningful guidance direction that promotes coherent global structure and naturally complements classifier-free guidance. On ImageNet 256×256256\times256, PerF-L achieves FID of 1.911.91, approaching 1.861.86 of JiT-H with only half the parameters, while PerF-H further achieves FID of 1.631.63 and 1.761.76 on ImageNet 256×256256\times256 and 512×512512\times512, respectively.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PixelDiT2: Representation-Grounded Pixel Diffusion Transformers

    Sep 21, 2026Yongsheng Yu, Wei Xiong, Yichen Sheng +2Contextual Grounding

  2. Advanced Pixel Diffusion Model with Guided Sparse Global Refinement

    Sep 1, 2026Weiyi You, Jinhua Zhang, Xingyu Zhou +3PixelloopPixels

  3. DuSPiT: Dual-Branch Sub-Patch Pixel Diffusion Transformer

    Jul 20, 2026Yunpeng Bai, Yossi Gandelsman, Michaël GharbiDiffusion TransformersMulti-Reference Image Generation