cs.AISep 29, 2026

DIET: Deletion-response Expert Trimming for Video Diffusion Transformers

Authors: Jiachang Zhang, Teng Hu, Bohao Feng, Songhang Shen, Wenqiang Wang, Hongqian Deng, Ran Yi

Organizations: Xi’an Jiao Tong University, China · Shanghai Jiao Tong University, China · Alibaba Token Hub, Alibaba Group, China · Alibaba Cloud Computing, China

Abstract

Video diffusion transformers (DiTs) increasingly adopt mixture-of-experts (MoE) architectures to reduce active computation, but their full expert storage remains costly. Existing one-shot pruning criteria mainly rely on static activation or routing statistics and cannot capture layer-level re-routing after expert deletion. We introduce DIET, a training-free expert pruning framework based on deletion responses. A single all-expert calibration pass records expert outputs and router states for matched conditional and unconditional tokens. Candidate deletions are then replayed from cached tensors, requiring no additional model forward passes. The resulting deletion-response signatures characterize each expert by the changes induced when it is removed. DIET selects retained experts by minimizing Overall Diversity Loss (ODL), which preserves directional coverage in signature space, and combines intra-layer local search with an inter-layer regression-guided budget search to allocate experts across layers. On LingBot-Video 30B-A3B, pruning 50% of experts (6,144 to 3,072) reduces the checkpoint from 57 GB to 30 GB and enables single-card deployment on a 48 GB GPU without fine-tuning. Under a fixed 284-case VBench protocol, the VBench Total increases from 0.7941 to 0.8115. Across tested retention budgets, DIET consistently outperforms competitive pruning baselines adapted from large language models.

Figures & tables

Appendix figures & tables29 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ScalingAttention: Discovering Intrinsic Sparse Attention Topology for Video Diffusion Transformers

    Jun 22, 2026Ruiliang Zhou, Xuecheng Wu, Kang He +6Dynamic Sparse AttentionDiffusion Transformers

  2. PARE: Pruning and Adaptive Routing for Efficient Video Generation

    May 26, 2026Yutong Wang, Yunke Wang, Tianfan Xue +4