cs.CVSep 29, 2026

Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

Authors: Yu Xu, Yuxin Zhang, Xiao Yang, Haotian Yang, Yizhi Wang, Xinwei Huang, Minxuan Lin, Angtian Wang, +2 more

Organizations: University of Chinese Academy of Sciences · ByteDance · Canva Research · University of Science and Technology Beijing

Abstract

Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing fragmentation and structural distortion. To address this, we propose SplitMoE, a split-role sparse architecture that breaks the shackles of uniformity. To accommodate the inherent semantic imbalance, we explicitly bifurcate the expert pool into semantic experts and generic experts, with semantic experts capturing high-level semantic abstraction and generic experts preserving residual visual information and flexible generative capacity. Leveraging prototype-guided routing and pull-push regularization, SplitMoE enables tokens to cluster naturally by semantic attributes rather than arbitrary balancing constraints. Extensive results show that under an equivalent activated-parameter budget, SplitMoE outperforms traditional load-balanced MoEs in convergence speed, routing coherence, and video generation quality across standard benchmarks. By revealing an emergent coarse-to-fine denoising logic, SplitMoE provides the community with a modality-aware scaling path, serving as a critical reference for building large-scale video world models.

Figures & tables

Explore similar work

CardsList
  1. Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE

    Jun 25, 2026Haoyou Deng, Keyu Yan, Chaojie Mao +4Localized Lora-MoeMixture-Of-Experts

  2. DecoMoE: Decoupling Visual Propagation and Expert Computation for Efficient Multimodal MoE Inference

    Sep 30, 2026Xudong Tan, Peng Ye, Ming Xie +4Mixture-Of-Expert InferenceMixture-Of-Experts