Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing fragmentation and structural distortion. To address this, we propose SplitMoE, a split-role sparse architecture that breaks the shackles of uniformity. To accommodate the inherent semantic imbalance, we explicitly bifurcate the expert pool into semantic experts and generic experts, with semantic experts capturing high-level semantic abstraction and generic experts preserving residual visual information and flexible generative capacity. Leveraging prototype-guided routing and pull-push regularization, SplitMoE enables tokens to cluster naturally by semantic attributes rather than arbitrary balancing constraints. Extensive results show that under an equivalent activated-parameter budget, SplitMoE outperforms traditional load-balanced MoEs in convergence speed, routing coherence, and video generation quality across standard benchmarks. By revealing an emergent coarse-to-fine denoising logic, SplitMoE provides the community with a modality-aware scaling path, serving as a critical reference for building large-scale video world models.
Figures & tables
Figure 1 : SplitMoE scales video diffusion models by introducing semantic-aligned expert partitioning, reducing spatiotemporal fragmentation, and generating higher-quality videos than dense fine-tuning and WAN with a standard MoE.
Figure 2 : Motivation for semantic-aware Video-MoE routing. Visual tokens exhibit much stronger cohesion than text tokens, making uniform MoE routing prone to fragmented assignments. Compared with standard MoE, SplitMoE routes tokens from different semantic regions to more distinguishable experts while assigning tokens within the same semantic region to more consistent expert groups.
Figure 3 : Pipeline of SplitMoE. SplitMoE splits each Video MoE layer into Semantic and Generic MoE branches, where VAE-prototype induced soft targets guide semantic routing while the generic branch preserves flexible residual modeling capacity. Learnable prototypes are optimized in the VAE feature space with pull-push regularization and a global token bank, encouraging semantic experts to form diverse and well-covered visual concepts.
Figure 4 : Qualitative comparison with baseline methods. Zoom in for better comparison. We provide more examples and comparisons with more baseline methods in the video demo.
VBench-2
T2V-CompBench
Model name
#Params.
Creativity ↑
Common Sense ↑
Control- lability ↑
Human Fidelity ↑
Physics ↑
Consist attr. ↑
Inter- action ↑
Nu- meracy ↑
Dense Wan2.2-FT
14B
49.57%
55.92%
30.37%
73.08%
54.66%
77.51%
61.08%
36.71%
Wan2.2-MoE
A14B
53.88%
57.15%
35.51%
73.49%
63.32%
78.82%
63.89%
38.55%
SplitMoE w/o PG
A14B
51.05%
63.22%
34.83%
74.04%
55.89%
78.46%
61.96%
38.57%
SplitMoE w/o Push
A14B
50.54%
62.75%
33.92%
73.60%
56.44%
77.29%
59.05%
37.15%
SplitMoE w/o Pull
A14B
52.23%
61.10%
35.67%
73.81%
56.83%
78.14%
61.15%
38.32%
Table 1 : Text-to-Video evaluation results on VBench-2 and T2V-CompBench, best scores in bold and second scores with underline . A14B denotes MoE with 14B activated parameters.
Method
Ours
Baseline MoE
Dense Model
Total Parameters
27B
27B
14B
Activated Parameters
14B
14B
14B
Training Speed
6.84s/it
6.69s/it
5.66s/it
Inference Speed
6.08s/it
6.12s/it
5.76s/it
Table 2 : Speed comparison.
Figure 5 : Ablation study analysis.
Figure 6 : Comparison of spatial routing and expert loads. (a) Baseline MoE scatters correlated patches across experts under strict capacity constraints, causing fragmented routing and poor spatiotemporal consistency. (b) SplitMoE routes coherent entities to Semantic Experts and residual textures to Generic Experts, achieving stable utilization without enforcing uniformity.
Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling diffusion models in visual generation. Recent advancements have focused on adaptively allocating computational resources across diverse tokens to improve efficiency and performance. However, we identify a routing assignment problem in existing diffusion MoE frameworks: the router fails to accurately allocate more computational resources to salient tokens. Our analysis attributes this failure to the router's reliance on noise-corrupted latent features throughout the denoising process. Such stochastic noise obscures the critical structural and textural information, thereby preventing the router from effectively distinguishing salient tokens. To address this, we propose SharpMoE, a post-training framework with a saliency-harnessing accurate routing mechanism, which utilizes clean latent features as a noise-free guidance signal for routing. By bypassing the noise-distorted inputs, SharpMoE provides the router with clear saliency guidance, enabling the identification of salient tokens even in high-noise stages. Furthermore, we introduce a trajectory routing loss to constrain the compute allocation throughout the multi-step denoising trajectory, ensuring precise resource allocation along the generation rollout. Extensive experiments demonstrate that SharpMoE serves as a versatile, plug-and-play solution that further enhances the pretrained, converged MoE models, achieving state-of-the-art performance in visual generation.
Haoyou Deng, Keyu Yan, Chaojie Mao +4
Key Laboratory of Image Processing and Intelligent Control, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology · Tongyi Lab, Alibaba Group
This paper systematically diagnoses the training failure modes of Token-Choice sparse Mixture-of-Experts (MoE) on video Diffusion Transformers. Starting from a pretrained dense model of about 5 billion parameters, we convert it into an MoE architecture following three laws: routed experts exactly clone the original FFN weights, shared experts are initialized to zero for verification and then to extremely small non-zero noise for actual training, while only the gating networks start from random initialization. Experiments reveal a hierarchy of five failure modes: (1) linear routers suffer global soft saturation with complete expert homogenization; (2) MLP routers introduce selective deadlock, where roughly one-third of layers degenerate into a single-expert mode that cannot be prevented by increasing the auxiliary loss; (3) cross-attention routers exhibit preliminary self-recovery, yet about nine layers remain stubbornly deadlocked; (4) deadlocked layers display a U-shaped distribution, concentrated in shallow visual processing layers and deep semantic integration layers; (5) bfloat16 mixed precision causes tiny weight updates to be truncated to zero by hardware. Based on routing decision time series over 65 million tokens across 5,000 training steps, we propose the Functional Redundancy Hypothesis: deadlock is a rational waiting strategy before the shared expert matures within the gate-shared expert-routed expert triadic system. This hypothesis is supported by the theory of functional redundancy in systems biology. On the engineering side, we summarize the Three Laws of dense-to-MoE conversion and provide a complete solution for the bfloat16 precision trap. We calibrate the current capability boundary of the Token-Choice paradigm and outline a three-step evolutionary roadmap from visual unification to a world model.
Multimodal mixture-of-experts (MoE) models combine sparse expert activation with visual-language capabilities, yet their inference remains costly because long visual-token sequences repeatedly incur attention, routing, dispatch, and expert-MLP computation. Existing methods typically compress either the token or expert dimension, leaving redundancy along the other. Our analysis reveals two complementary regularities: the depth required for visual propagation varies across inputs, while text-token routing exhibits concentrated and recurrent expert-importance patterns. Based on these observations, we propose DecoMoE, a two-dimensional structured compression framework that decouples visual propagation from expert computation. The Sample-Adaptive Visual Boundary (SAVB) predicts an input-dependent visual-exit layer at which the visual-token block is removed. The Routing-Calibrated Expert Prefix (RCEP) reorders experts offline using text-token routed mass and, from this predicted exit layer onward, retains at each MoE layer the shortest contiguous prefix covering a target routed-mass fraction. We evaluate DecoMoE on Qwen3-VL-MoE and InternVL3.5-30B-A3B across six benchmarks. On Qwen3-VL-MoE, DecoMoE retains 97.91% of dense-baseline performance while reducing computation from 27.06 to 16.73 TFLOPs and latency from 0.44 to 0.26 seconds, yielding a 1.69x speedup. Code will be available at https://github.com/ShawnTan86/DecoMoE.
Xudong Tan, Peng Ye, Ming Xie +4
Fudan University · The Chinese University of Hong Kong