Mixture-of-Depths (MoD) enables conditional computation across Transformer depth by routing only a subset of tokens through selected layers, but its original one-sparse--one-dense alternation tightly couples total capacity to active capacity and limits sparse-depth scaling. We introduce X-MoD, a scalable sparse-depth architecture that decouples token sparsity from anchor stride, allowing total parameter count to grow while keeping active-equivalent capacity nearly fixed. To make deep sparse routing trainable, X-MoD combines dense anchors with variance-scaled layer-wise gating and depth-wise token balancing. To make this regime analyzable and usable, we formulate sparse-depth routing as a conditional architecture-design problem: given compute, context length, and active-equivalent backbone size, how should the routing configuration be chosen? We develop a practical scaling-law framework by fitting X-MoD relative to FLOP-matched dense baselines, yielding an interpretable law that decomposes performance into sparse-capacity gain, sparse-context correction, and anchor-stride interaction. The law predicts validation loss across routing configurations and reveals how context length, model scale, and anchor stride shape sparse-depth performance. We validate the architecture and law through pretraining sweeps, held-out scaling-law prediction, ablations, downstream evaluations, and comparisons with Dense, MoD, and representative MoE baselines.
Figures & tables
Figure 1: MoD vs. X-MoD. X-MoD decouples token sparsity K from anchor stride A , enabling a substantially larger gap between total and active-equivalent parameter counts.
Figure 2: Empirical regularities motivating the X-MoD scaling law. (a, b) Loss is U-shaped in K , and the optimum shifts with context length L , also visible under sparse context Ls=L/K . (c) Varying Nact has weaker effect on optimal K . (d) Anchor stride A interacts with K . Sweep settings are given in Appendix B.4 .
Figure 3: Predicted vs. true validation loss. (a) In-range fit (circles) and out-of-range predictions without refitting (squares). (b–d) Leave-one-group-out predictions by L , A and Nact , respectively.
Nact
Model
N
ϕ/ϕD↓
Loss ↓
Hella. ↑
ARC-e ↑
ARC-c ↑
PIQA ↑
LAMB. ↑
BoolQ ↑
Avg. ↑
556M
Dense
0.54B
1.00
2.773
30.42
51.93
23.21
63.80
18.47
60.52
41.39
MoD
0.87B
0.92
2.692
34.00
58.75
25.94
66.43
21.33
51.99
43.07
MoE 8:64
2.46B
1.00
2.564
36.62
62.25
27.30
68.99
26.37
60.95
47.08
X-MoD A4K8
2.64B
0.47
2.549
38.30
62.08
27.56
68.34
29.19
58.62
47.35
MoE 8:128
4.58B
1.00
2.528
38.57
64.98
30.38
70.18
27.25
56.30
47.94
X-MoD A3K16
4.79B
0.46
2.503
38.00
64.31
28.41
70.13
30.49
60.49
48.64
Table 1: Main results under matched active-equivalent size and training compute. N denotes total parameters, and ϕ/ϕD denotes per-token training FLOPs normalized by the dense model with the same Nact and sequence length. The best results are highlighted in bold . Detailed compute budgets and full model configurations are reported in Appendix B .
Variant
Total N
ϕ/ϕD
Loss ↓
Hella. ↑
ARC-e ↑
ARC-c ↑
PIQA ↑
LAMB. ↑
BoolQ ↑
Avg. ↑
Dense
936M
1.00
2.635
35.18
59.85
26.71
67.52
24.16
59.85
45.54
MoD
1.48B
0.92
2.570
36.67
62.04
26.54
68.72
26.14
60.89
46.83
Routed-FFN A4K8
3.26B
1.00
2.508
37.97
64.27
28.92
70.29
29.79
60.95
48.70
Routed-FFN A3K16
5.68B
1.00
2.482
40.71
65.61
30.38
69.15
29.89
60.37
49.35
Full X-MoD A4K8
4.52B
0.51
2.427
41.26
68.69
33.45
71.22
33.55
60.43
51.43
Selected-Q / full-KV A4K8
4.52B
1.09
2.564
38.44
61.66
27.65
68.01
28.90
57.71
47.06
Table 2: Mechanism controls of X-MoD. All variants use the 936M reference backbone. Metrics follow Table 1 ; configurations are given in Appendix B.6 .
Variant
Loss ↓
Δ Loss
Hella. ↑
ARC-e ↑
ARC-c ↑
PIQA ↑
LAMB. ↑
BoolQ ↑
Avg. ↑
Dense reference
2.635
+0.159
35.18
59.85
26.71
67.52
24.16
59.85
45.54
Full X-MoD A1K16
2.476
0.000
40.64
66.08
30.80
71.27
30.58
61.31
50.12
w/o dense anchors
2.579
+0.103
36.35
62.63
28.33
69.53
26.59
58.96
47.06
w/o gated residual scaling
2.492
+0.016
40.62
65.19
29.27
69.31
29.73
60.58
49.12
w/o token-choice bias
2.657
+0.181
34.59
59.85
26.54
67.52
22.01
59.11
44.94
w/o dense prefix
2.487
+0.011
40.41
65.19
30.20
69.70
29.69
61.01
49.37
Table 3: Ablations of X-MoD stabilization mechanisms. All variants are compared under the same Nact=936 M. Δ Loss is measured relative to Full X-MoD; task scores follow Table 1 .
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Nact
Dim
Head dim
Layers
Heads
KV heads
FFN dim
165M
512
64
28
16
8
1536
298M
768
64
28
16
8
2304
556M
1024
128
28
16
8
3072
936M
1536
128
28
16
8
3840
1.65B
2048
128
28
16
8
6144
Appendix
Table 4: Backbone model configurations. All models use GQA attention and standard feed-forward layers.
Hyperparameter
Value
Peak learning rate
3.2×10−3
Momentum coefficient β
0.95
Newton–Schulz steps
5
Nesterov momentum
True
RMS match
0.2
AdamW betas
(0.9,0.95)
Appendix
Table 5: Muon optimizer configuration selected in preliminary dense-baseline sweeps and used in the main runs.
Nact
C
Model
Layers
Experts
de
N
FLOPs/tok.
Exec.
556M
1020
Dense
28
–
–
0.54B
8.4×109
DDP
MoD
49
–
–
0.87B
7.7×109
DDP
MoE 8:64
28
6:64+2
384
2.46B
8.4×109
DDP
X-MoD A4K8
161
–
–
2.64B
3.9×109
DDP
MoE 8:128
28
6:128+2
384
4.58B
8.4×109
DDP
X-MoD A3K16
298
–
–
4.79B
3.9×109
DDP
Appendix
Table 6: Detailed model configurations used in the main comparison. Nact denotes active-equivalent parameters, N denotes total parameters and C denotes total training FLOPs.
a, Routed computation and attention context
Model
Routed computation
KV context
Dense
None
Full
MoD
Attention + FFN
Selected subset
Routed-FFN A4K8
FFN only
Full every K refinements
Routed-FFN A3K16
FFN only
Full every K refinements
Selected-Q/full-KV A4K8
Q/O + FFN
Full sequence
Appendix
Table 7: Mechanism-control configurations at the 936M reference scale. D/DD=(ϕ/ϕD)−1 denotes token exposure relative to Dense.
Table 8: Fitted parameters of the practical X-MoD scaling law. For ease of reference, Eq. ( 14 ) is reproduced above the estimates.
Figure 4: Pilot comparison of X-MoD and MoE configurations. Validation loss versus total parameters at Nact=165M , L=32k and C=3×1019 FLOPs. Circles denote X-MoD and diamonds denote MoE; colour indicates K , and marker area increases with total parameters. Green dashed lines mark the selected MoE parameter budgets; black outlines identify the selected MoE–X-MoD pairs. The plot labels 8o64s2 and 8o128s2 correspond to MoE 8:64 and 8:128.
Variant
Total N
N/Nact
Total layers
ϕ/ϕD
Val. loss ↓
Deep Dense
256M
1.00
52
1.00
3.084
Deep X-MoD A1K4
552M
2.16
124
0.67
3.047
Deep X-MoD A1K8
939M
3.67
220
0.62
3.034
Deep X-MoD A1K12
1.29B
5.04
316
0.60
3.023
Deep X-MoD A1K16
1.67B
6.52
412
0.59
3.011
Appendix
Table 9: Total-depth scaling of X-MoD. All variants use the width of the 165M backbone, with matched active-equivalent capacity and training compute C=4×1019 .
Repeated identical subset
Independent random subsets
Threshold routing
Threshold / top- k
Model
Coverage
Jaccard
Coverage
Jaccard
Coverage
Jaccard
Agreement
X-MoD A1K16
6.250
100.0
64.39
3.226
52.24
14.51
98.63
A1K16 w/o token-choice bias
6.250
100.0
64.39
3.226
58.59
80.66
–
X-MoD A4K8
12.50
100.0
98.61
6.667
99.99
1.947
96.62
X-MoD A3K16
6.250
100.0
95.49
3.226
97.50
2.940
98.25
Appendix
Table 10: Routing coverage, consecutive-layer overlap and mask agreement. All entries are percentages at the 936M reference scale; a dash denotes an unmeasured value.
Figure 5: Token-wise sparse-update distributions under threshold routing. (a) A1K16 with and without token-choice bias (16-layer intervals). (b) A4K8 (complete 32-layer intervals). (c) A3K16 (48-layer intervals). Grey dashed and green dot-dashed curves show repeated-subset and independent-subset references. Curves interpolate discrete probabilities; integer heights, not filled areas, represent probability. Models and evaluation conditions follow Table 10 .
a, Model scale, counted compute and peak training memory
Model
Layers
N
FLOPs/token ↓
Peak memory (GB/GPU) ↓
Dense
28
936M
9.0G
32.08
MoD
49
1.48B
8.3G
46.62
MoE 8:64
28
4.51B
9.0G
55.46
X-MoD A4K8
161
4.52B
4.6G
55.72
MoE 8:128
28
8.48B
9.0G
71.96
Appendix
Table 11: Measured systems efficiency at the 936M active-equivalent scale. Speedups are relative to the adjacent, approximately total-parameter-matched MoE baseline.
While Mixture-of-Experts (MoE) scales model capacity without proportionally increasing computation, its massive total parameter footprint creates significant storage and memory-access bottlenecks, which hinder efficient end-side deployment that simultaneously requires high performance, low computational cost, and small storage overhead. To achieve these properties, we present DECO, a sparse MoE architecture designed to match the performance of dense Transformers under identical total parameter budgets and training tokens. DECO utilizes the differentiable and flexible ReLU-based routing enhanced by learnable expert-wise scaling, which adaptively balances the contributions of routed and shared experts. Furthermore, we introduce NormSiLU, an activation function that normalizes inputs prior to SiLU operators, producing a more stable trend of routed-expert activation ratio and a higher intrinsic sparsity level. We also identify an empirical advantage in using non-gated MLP experts with ReLU-based routing, indicating the possibility of MoE architecture simplification. Experiments demonstrate that DECO, activating only 20% of routed experts, matches dense performance and outperforms established MoE baselines. Our specialized acceleration kernel delivers a 2.93× speedup on Jetson AGX Orin compared with dense inference. Code and checkpoints are available at https://github.com/thunlp/DECO.
Chenyang Song, Weilin Zhao, Xu Han +3
Dept. of Comp. Sci. & Tech., Institute for AI, Tsinghua University, Beijing, China
In large-scale pretraining, the algorithm, architecture, and systems decisions are conventionally made in disconnected stages. A scaling law stage selects an architecture and training recipe, optimizing loss under compute constraints, and a separate systems stage then optimizes the implementation for hardware efficiency. In this work, we develop MOSAIC, which formulates model architecture and systems co-design as an optimization problem. MOSAIC couples a predictive scaling law with a calibrated performance model that estimates Model FLOPs Utilization (MFU), communication cost, memory footprint, and the best parallel layout. We instantiate the framework for sparse Mixture-of-Experts (MoE) language models, where expert count, routing sparsity, and other MoE layer dimensions affect both the loss and systems efficiency. We fit a scaling law on sparse MoE models trained on text data, whose scaling dimensions include the sparsity factor, which is the fraction of model parameters inactive per token in a forward pass. The scaling law sweeps in our work span active parameters from 104 million to 2.7 billion and total model sizes reaching 79 billion parameters. We show that, within the calibrated sparsity range, an efficiency-agnostic model-FLOPs budget admits no interior optimal sparsity. The fitted loss decreases monotonically with sparser models and the compute optimum lies at the upper boundary of the data support. An optimal sparsity in MoE models instead emerges under the cluster's systems constraints, as captured by MOSAIC. Our results argue for a shift towards unified architecture and systems co-design for frontier language model training.
Standard Mixture-of-Experts (MoE) transformers route tokens to expert subnetworks within each layer, but the layer structure itself remains monolithic. We introduce Mixture of Layers (MoL), which replaces full-width transformer blocks (d_model) with K parallel thin blocks at reduced dimensionality (d_thin << d_model), connected via learned down/up projections and composed via top-k block routing. Scaling sparse block routing to many blocks creates an attention coverage problem, as each block sees fewer tokens. We address this by introducing hybrid attention, which pairs one shared softmax block for global context with Gated DeltaNet linear attention in routed blocks.
Ivan Ternovtsii, Yurii Bilak
Department of Software Systems, Uzhhorod National University Narodna sq. 3, Uzhhorod, Ukraine, 88000