Video diffusion transformers (DiTs) increasingly adopt mixture-of-experts (MoE) architectures to reduce active computation, but their full expert storage remains costly. Existing one-shot pruning criteria mainly rely on static activation or routing statistics and cannot capture layer-level re-routing after expert deletion. We introduce DIET, a training-free expert pruning framework based on deletion responses. A single all-expert calibration pass records expert outputs and router states for matched conditional and unconditional tokens. Candidate deletions are then replayed from cached tensors, requiring no additional model forward passes. The resulting deletion-response signatures characterize each expert by the changes induced when it is removed. DIET selects retained experts by minimizing Overall Diversity Loss (ODL), which preserves directional coverage in signature space, and combines intra-layer local search with an inter-layer regression-guided budget search to allocate experts across layers. On LingBot-Video 30B-A3B, pruning 50% of experts (6,144 to 3,072) reduces the checkpoint from 57 GB to 30 GB and enables single-card deployment on a 48 GB GPU without fine-tuning. Under a fixed 284-case VBench protocol, the VBench Total increases from 0.7941 to 0.8115. Across tested retention budgets, DIET consistently outperforms competitive pruning baselines adapted from large language models.
Figures & tables
Figure 1: (a) Grouped routing in LingBot-Video: each token activates eight of the 128 experts of a layer through grouped top- 2 routing, and a single all-expert calibration records every expert output and the router state. (b) Deleting an expert and replaying the grouped top- k reproduces the deletion at the captured states from cached tensors, without extra expert forwards, and the case-level responses concatenate into a deletion-response signature . (c) The overall diversity loss (ODL) matches every deleted expert to its nearest retained expert in signature space and sums the matched distances, and the retained set is optimized under a given per-layer budget. (d) The regression-guided search allocates the retention ratio across layers: solid curves, our allocations at 20% , 50% and 80% retention; dashed, uniform references.
Retention Budget
Total ↑
Quality ↑
Semantic ↑
Unpruned (100%)
0.7941
0.8125
0.7204
20% ( 1,229 experts)
0.7596
0.7958
0.6145
50% ( 3,072 experts)
0.8115
0.8325
0.7278
80% ( 4,915 experts)
0.8061
0.8331
0.6980
Table 1: Performance across retention budgets on official VBench ( 284 -case paired protocol). Expert budgets correspond to 1,229 , 3,072 , and 4,915 retained experts out of 6,144 .
Figure 2: Primary evaluation and routing dynamics. (a) VBench Total across candidate retention budgets on the 142 -case budget-search set (dashed line denotes the dense baseline). (b) Number of original routed experts retained per token (out of 8). (c) Reallocated gating weight assigned to promoted surviving experts (averaged over 120 calibration cases, 48 layers). Counter-intuitively, DIET induces the highest routing perturbation while achieving the highest benchmark scores, demonstrating that routing preservation is a suboptimal objective.
20% Retention
50% Retention
80% Retention
Method
Total
Qual.
Sem.
Total
Qual.
Sem.
Total
Qual.
Sem.
Uniform Layer Budget
DIET (Ours)
0.7370
0.7742
0.5883
0.7996
0.8184
0.7243
0.8052
0.8311
0.7013
REAP
0.6756
0.7645
0.3204
0.7807
0.8033
0.6901
0.7943
0.8171
0.7028
SHAPE
0.6647
0.7507
0.3209
0.7374
0.7662
0.6223
0.7850
0.8081
0.6929
HTS
0.6909
0.7801
0.3342
0.7921
0.8252
0.6595
0.8025
0.8274
0.7027
Table 2: Official VBench comparative evaluation ( 284 -case protocol, identical sampling seeds and evaluation pipeline). Retention budgets correspond to 1,229 ( 20% ), 3,072 ( 50% ), and 4,915 ( 80% ) surviving experts.
Figure 3: Per-dimension deviation from the unpruned baseline at 50% retention.
Component
Configuration
Total ↑
Space
Router Frequency
0.7485
Router Score
0.7813
Output Space
0.7883
Objective
Summed-Response Norm
0.7571
PCA Energy ( r=32 )
0.7849
D-Optimal Volume
0.8007
Table 3: Ablation analysis on representation spaces, objectives, and optimization strategies at 50% retention.
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Shared 142
Held-Out 142
Full Model (Unpruned)
0.8023
0.7858
Router Frequency
0.7598
0.7367
REAP ( Lasby et al., 2026 )
0.7784
0.7745
SHAPE ( Zhang, 2026 )
0.7363
0.7569
HTS ( Liu et al., 2026b )
0.7785
0.7858
TB-Coverage ( Zeng et al., 2026 )
0.7788
0.7637
Appendix
Table A1: Comparative performance partitioned across the 284 -case protocol: the 142 cases overlapping the search prompt distribution and the 142 held-out cases unseen during search. Evaluation adheres to official VBench normalization ( n=139 – 142 completed pairs per cell).
Retention
Retained
Total
Kept Range
20%
1,229
0.7710
16–116
30%
1,843
0.8016
18–124
40%
2,458
0.8015
16–128
50%
3,072
0.8229
52–79
60%
3,686
0.8103
64–94
70%
4,301
0.8039
33–128
Appendix
Table A2: Exploration sweep on the 142 -case budget-search set; Total aggregates all completed cases of each run (selection-time scores).
Figure A1: Layer-wise retained expert counts across exploration budgets. Colors denote retaining counts out of 128 candidate experts. The selected 50% mask allocates capacity in a balanced band ( 52 – 79 experts per layer), whereas unconstrained exploration budgets assign as few as 16 experts to resilient layers.
Figure A2: Structural characteristics of the selected 50% mask. (a) Full 48×128 retaining allocation pattern (dashed lines indicate group boundaries). (b) Retained counts per group per layer. (c) Total retained experts per layer. Every group retains at least 10 experts and each layer preserves at least 52 , strictly satisfying grouped top- 8 routing feasibility.
Figure A3: Qualitative comparisons across visual appearance, background, and color fidelity dimensions. Each panel depicts representative generations under identical initializations for the dense baseline (top) and DIET at 50% retention (bottom).
Figure A4: Qualitative comparisons across dynamic degree, human action, imaging quality, and motion smoothness (protocol matches Figure A3 ).
Figure A5: Qualitative comparisons across multiple objects, object class, overall consistency, and scene composition (protocol matches Figure A3 ).
Figure A6: Qualitative comparisons across spatial relationship, subject consistency, temporal flickering, and temporal style (protocol matches Figure A3 ).
Figure A7: All-arm frame-level comparisons on two representative cases: (a) a ballroom scene (“ballroom”) and (b) an aesthetic-quality prompt (“An astronaut feeding ducks on a sunny afternoon, reflection from the water.”). Each panel shows DIET at 50% retention, the unpruned baseline, and the five ported criteria (named as in Table 2 ) as rows, with five evenly spaced frames of the same generation across columns. In (a) the scene composition is preserved across all arms, while in (b) the astronaut’s suit collapses into a featureless mass under SHAPE and Routing Frequency; DIET remains visually consistent with the unpruned model in both panels.
Figure A8: Layer-wise routing concentration profiles across conditional and unconditional branches. (a) Gini coefficients and normalized entropy across depth. (b) Token share absorbed by the top-decile experts. Routing remains well-dispersed across all layers, demonstrating the absence of inactive capacity.
Figure A9: Stability across calibration sample sizes. (a) Retained expert mask overlap between C -case solutions and the 120 -case reference (dashed line indicates multi-seed solver variance of 0.912 ). (b) Relative ODL objective value evaluated against the full calibration set. Reducing calibration size preserves functional objective quality within 2.7% .
Selection Objective (Matched Budget)
Total ODL ↓
High- cos Deletions ↑
VBench Total ↑
DIET (ODL, Response Space)
2561.2
135
0.8115
D-Optimal (Volume Maximization)
2648.1
62
0.8007
PCA Energy ( r=32 )
2732.9
5
0.7849
PCA Energy ( r=8 )
2755.8
3
0.7837
ODL (Isolated Output Space)
2676.1
74
0.7883
ODL (Router Score Space)
2692.6
77
0.7813
Appendix
Table A3: End-to-end comparison of candidate selection objectives under identical layer budgets ( 3,072 retained experts). High- cos deletions denote pruned experts whose nearest retained expert exhibits cos>0.3 . Low-rank PCA truncation and space-ablated variants degrade performance by 2.32 – 3.02 points.
Selection Rule
Total ODL
Mean Loss
Deletions
Deletions
Mean ∥se∥
∑eδe↓
Per Expert ↓
( cos>0.3 ) ↑
( cos>0.5 ) ↑
of Excised
DIET ( 50% )
2561.2
0.834
135
27
0.706
HTS ( Liu et al., 2026b )
2760.4
0.899
23
2
0.521
REAP ( Lasby et al., 2026 )
2763.7
0.900
22
2
0.484
SHAPE ( Zhang, 2026 )
2767.4
0.901
26
1
0.468
Router Frequency
2769.7
0.902
21
1
0.463
Appendix
Table A4: Deletion profiles across pruning heuristics under matched layer budgets ( 3,072 retained experts, 3,072 deletions). Loss denotes δe=1−cos(se,sf(e)) . DIET specifically targets functionally covered experts ( 135 with cos>0.3 ) rather than indiscriminately pruning low-norm units.
Figure A10: Maximum pairwise signature cosine within the retained expert sets across four representative layers. DIET consistently maintains the lowest internal redundancy, preserving diverse directional coverage.
Figure A11: Geometric properties of deletion-response signatures at layer 24 . (a) Pairwise signature cosine matrix. (b) Distribution of pairwise cosines (logarithmic scale), demonstrating sharp concentration near zero with an extended positive tail. (c) Cumulative singular value energy, exhibiting absence of low-rank concentration.
Figure A12: Layer-wise geometric profile across all 48 transformer blocks. (a) Singular spectrum participation ratios. (b) Pairwise cosine quantiles and maxima. (c) Density of high-cosine pairs ( cos>0.3 ). (d) Mean nearest-neighbor cosine.
Figure A13: Empirical correlation between layer-wise retained expert counts and standardized VBench dimension scores across search trajectories, highlighting functional division of labor across network depth.
Method
Pruning Metric
Source Domain
Reference
Router Frequency
Token dispatch counts
LLM MoE
Standard Baseline
REAP
Gating weight × activation norm
LLM MoE
Lasby et al. (2026)
SHAPE
Coalition Shapley values
LLM MoE
Zhang (2026)
HTS
Activation norm formulations
LLM MoE
Liu et al. (2026b)
TB-Coverage
Round-robin register protection
LLM MoE
Zeng et al. (2026)
NAEE-Style Recon
Replayed layer-output reconstruction
LLM MoE
Lu et al. (2024)
Appendix
Table A5: Summary of evaluated MoE pruning baselines adapted to video DiT architectures ( Lasby et al., 2026 ; Zhang, 2026 ; Liu et al., 2026b ; Zeng et al., 2026 ; Lu et al., 2024 ) . All methods utilize identical calibration tensors, evaluation protocols, and budget allocations.
Dimension
Full model
Router freq.
REAP
SHAPE
HTS
TB-Coverage
NAEE-Style Recon
Ours
Aesthetic qual.
0.554
0.503
0.555
0.485
0.580
0.528
0.540
0.598
Appearance style
0.806
0.789
0.785
0.773
0.802
0.812
0.817
0.813
Background cons.
0.937
0.929
0.953
0.920
0.955
0.953
0.947
0.954
Color
0.845
0.868
0.675
0.768
0.738
0.794
0.905
0.809
Dynamic degree
0.733
0.467
0.600
0.600
0.733
0.600
0.733
0.800
Human action
1.000
0.905
0.952
0.905
0.905
0.952
0.952
1.000
Appendix
Table A6: Dimension-wise VBench evaluation on the full 284 -case protocol under matched 50% budgets ( 3,072 retained experts).
Dimension
Full model
Router freq.
REAP
SHAPE
HTS
TB-Coverage
NAEE-Style Recon
Ours
Aesthetic qual.
0.554
0.336
0.423
0.355
0.415
0.423
0.436
0.535
Appearance style
0.806
0.762
0.762
0.767
0.787
0.757
0.771
0.804
Background cons.
0.937
0.932
0.967
0.931
0.957
0.958
0.925
0.935
Color
0.845
0.818
0.662
0.700
0.615
0.762
0.764
0.769
Dynamic degree
0.733
0.667
0.267
0.533
0.533
0.600
0.467
0.400
Human action
1.000
0.286
0.524
0.381
0.762
0.476
0.571
0.857
Appendix
Table A7: Dimension-wise VBench evaluation on the full 284 -case protocol at 20% retention ( 1,229 retained experts).
Dimension
Full model
Router freq.
REAP
SHAPE
HTS
TB-Coverage
NAEE-Style Recon
Ours
Aesthetic qual.
0.554
0.585
0.576
0.577
0.572
0.589
0.579
0.586
Appearance style
0.806
0.814
0.799
0.806
0.795
0.811
0.814
0.800
Background cons.
0.937
0.934
0.958
0.944
0.953
0.957
0.954
0.959
Color
0.845
0.797
0.769
0.808
0.783
0.794
0.861
0.790
Dynamic degree
0.733
0.800
0.867
0.800
0.867
0.867
0.867
0.867
Human action
1.000
1.000
1.000
1.000
1.000
1.000
0.952
1.000
Appendix
Table A8: Dimension-wise VBench evaluation on the full 284 -case protocol at 80% retention ( 4,915 retained experts).
Figure A14: Cross-resolution evaluation. (a) Total at native 480 p for the unpruned model and the pruned mask, on the 71 matched cases of the resolution screen. (b) Paired dimension-wise differentials at native 480 p (pruned mask minus unpruned, 95% CI).
Figure A15: Representative native 480 p generations across positive, neutral, and sensitive dimensions under matched initializations. Panels compare unpruned baselines (top row) against DIET at 50% retention (bottom row).
Metric
Full Model
DIET ( 50% )
Routed Expert Count
6,144
3,072
Expert Bank Parameters
29.0B
14.5B
Checkpoint Disk Footprint ( bfloat16 )
57 GB
30 GB
Required GPUs per Replica ( 48 GB VRAM)
2
1
Total Replica VRAM Allocation
93.2 GiB ( 2×46.6 )
44.7 GiB
Peak Memory ( 240×416 , Batch Size 2)
46.6 GiB
44.7 GiB
Appendix
Table A9: Physical deployment and memory footprint comparison between the unpruned model and DIET at 50% retention. Measurements reflect empirical single-replica execution. Active FLOPs per token remain strictly identical.
Figure A16: Breakdown of additive superposition under joint multi-expert deletions. Cosine similarity between true joint response perturbations and constituent single-deletion sums degrades sharply with deletion cardinality.
Figure A17: Paired per-case bootstrap differentials ( 95% percentile intervals) comparing DIET against the dense baseline and ported MoE pruning heuristics at 50% retention.
Figure A18: Dimension-wise paired differentials ( 95% bootstrap intervals) comparing DIET against (a) the unpruned baseline and (b) ported LLM pruning baselines, including the reconstruction-criterion port NAEE-Style Recon. While differentials against the unpruned baseline exhibit conservative preservation across dimensions, comparisons against ported baselines demonstrate decisive semantic advantages.
Figure A19: Dimension-wise paired differentials at (top) 20% and (bottom) 80% retention, with the same conventions as Figure A18 .
Figure A20: Per-dimension deviation from the unpruned baseline at (left) 20% and (right) 80% retention; Figure 3 shows the 50% version.