cs.LGSep 28, 2026

X-MoD: Practical Scaling Laws for Sparse-Depth Routing Beyond Mixture-of-Depths

Authors: Bowen Dong, Yilong Fan, Tengyu Pan, Yike Zhang, Zhenyu Li, Zijian Zhang, Xuewei Li, Mei Yu, +1 more

Organizations: Tsinghua University · Tianjin University

Abstract

Mixture-of-Depths (MoD) enables conditional computation across Transformer depth by routing only a subset of tokens through selected layers, but its original one-sparse--one-dense alternation tightly couples total capacity to active capacity and limits sparse-depth scaling. We introduce X-MoD, a scalable sparse-depth architecture that decouples token sparsity from anchor stride, allowing total parameter count to grow while keeping active-equivalent capacity nearly fixed. To make deep sparse routing trainable, X-MoD combines dense anchors with variance-scaled layer-wise gating and depth-wise token balancing. To make this regime analyzable and usable, we formulate sparse-depth routing as a conditional architecture-design problem: given compute, context length, and active-equivalent backbone size, how should the routing configuration be chosen? We develop a practical scaling-law framework by fitting X-MoD relative to FLOP-matched dense baselines, yielding an interpretable law that decomposes performance into sparse-capacity gain, sparse-context correction, and anchor-stride interaction. The law predicts validation loss across routing configurations and reveals how context length, model scale, and anchor stride shape sparse-depth performance. We validate the architecture and law through pretraining sweeps, held-out scaling-law prediction, ablations, downstream evaluations, and comparisons with Dense, MoD, and representative MoE baselines.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices

    May 11, 2026Chenyang Song, Weilin Zhao, Xu Han +3Mixture-Of-Experts ArchitecturesMixture-Of-Experts

  2. Mixture of Layers with Hybrid Attention

    May 10, 2026Ivan Ternovtsii, Yurii BilakTransformer AttentionMixture-Of-Experts