Mixture-of-Experts (MoE) Diffusion Language Models (DLMs) offer flexible parallel decoding and increased model capacity, but their large number of expert parameters incurs substantial computation and storage costs. Existing low-rank MoE compression methods largely rely on static factorization and fixed rank allocation, which overlook the distinctive properties of MoE DLMs. Specifically, we identify two properties: cross-mode non-uniform redundancy, where parameter redundancy and sensitivity to rank truncation vary across the input, output, and expert modes, and token-wise utilization variation, where hot and cold tokens exhibit distinct spectral characteristics and expert activation patterns. To address these challenges, we propose ITC-MoE, an Importance-guided Token-aware Compression framework for MoE DLMs. ITC-MoE consists of two complementary components. First, Importance-guided Adaptive Tucker Compression (IATC) incorporates activation and gradient importance into expert weight transformation, jointly factorizes expert weights across multiple modes, and adaptively allocates ranks under a fixed parameter budget. Second, Token-aware Compensation and Routing (TCR) applies lightweight low-rank compensation to compression-sensitive hot tokens and restricts the candidate expert set for cold tokens with concentrated routing patterns. By jointly adapting compression capacity and inference execution to both parameter redundancy and token-wise variation, ITC-MoE substantially reduces the computation and storage costs of MoE DLMs while preserving their generation quality. For example, on SDAR-30B-A3B-Chat-b32, ITC-MoE maintains an accuracy of 96.33% on MultiArith under a 30% compression budget, while achieving up to a 7.22x end-to-end speedup. The code is publicly available at https://github.com/lianjunl13-sudo/ITC-MoE.
Figures & tables
Figure 1 : Overview of the ITC-MoE framework.
Figure 2 : Observations. Results on SDAR-30B-A3B-Chat-b32 ( Cheng et al., 2026 ) . Panels (a)-(d) use the MBPP benchmark, while (e) and (f) use the GSM8K benchmark. (a,b) Expert weight spectra. Cumulative spectral energy of the layerwise expert-weight tensor along input- and output-feature modes, respectively. (c) Expert representation similarity. Pairwise linear CKA between expert output representations within a layer, where higher values indicate higher similarity. “mean off-diag.” is the mean off-diagonal CKA. (d) Mode sensitivity. Relative MoE output reconstruction error vs. dense reference when each Tucker mode is truncated independently with the other two at full rank. (e) Token-dependent rank. Ranks retaining 95% energy for hot and cold tokens across MoE layers. (f) Token-dependent expert patterns. Distribution of experts activated by cold and hot tokens under Top-8 routing.
M
Cold recall
Hot recall
16
0.8975
0.6930
32
0.9627
0.8186
48
0.9795
0.8834
64
0.9890
0.9257
Table 1 : Comparison of Top-8 recall for different candidate set sizes on cold and hot tokens. Top-8 recall is defined as the fraction of tokens whose ground-truth Top-8 expert set is fully contained in the candidate set. Evaluation uses the SDAR-30B-A3B-Chat-b32 model with held-out GSM8K and MBPP calibration prompts, with 128 experts and Top-8 routing per token.
Benchmark
Method
10% Ratio
20% Ratio
30% Ratio
Score ↑
Speedup ↑
Score ↑
Speedup ↑
Score ↑
Speedup ↑
MBPP
Original
65.76
1.00×
65.76
1.00×
65.76
1.00×
D 2 -MoE
56.42
1.26×
52.53
1.14×
34.24
0.80×
TD-MoE
58.75
1.34×
51.75
1.31×
12.45
1.35×
ITC-MoE
59.92
5.68×
50.58
5.68×
47.47
6.61×
GSM8K
Original
90.60
1.00×
90.60
1.00×
90.60
1.00×
Table 2 : Comparison of D 2 -MoE ( Gu et al., 2025 ) , TD-MoE ( XU et al., 2026 ) , and our ITC-MoE on SDAR-30B-A3B-Chat-b32 Cheng et al. (2026) under different compression ratios. The orange-signed numbers indicate that the corresponding speedups are below 1.00× .
Benchmark
Method
SDAR
LLaDA2.0-mini
Score ↑
Speedup ↑
Score ↑
Speedup ↑
MBPP
Original
65.76
1.00×
80.54
1.00×
TEAM
65.76
2.08×
80.16
1.56×
ITC-MoE
65.57
4.61×
80.93
4.53×
GSM8K
Original
90.60
1.00×
91.89
1.00×
TEAM
90.30
1.83×
92.80
1.72×
Table 3: Comparison of ITC-MoE and TEAM across different backbone models, with model compression applied to ITC-MoE.
Figure 3 : Hot-token threshold sensitivity and effect of Importance-Weighted Transformation.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Layers
MoE Layers
Routed Experts
Experts/Token
SDAR-30B-A3B-Chat-b32
48
48
128
8
LLaDA2.0-mini
20
19
256
8
Appendix
Table 6: Architectural configurations of the evaluated MoE-based diffusion language models.