Mixture-of-Depths (MoD) enables conditional computation across Transformer depth by routing only a subset of tokens through selected layers, but its original one-sparse--one-dense alternation tightly couples total capacity to active capacity and limits sparse-depth scaling. We introduce X-MoD, a scalable sparse-depth architecture that decouples token sparsity from anchor stride, allowing total parameter count to grow while keeping active-equivalent capacity nearly fixed. To make deep sparse routing trainable, X-MoD combines dense anchors with variance-scaled layer-wise gating and depth-wise token balancing. To make this regime analyzable and usable, we formulate sparse-depth routing as a conditional architecture-design problem: given compute, context length, and active-equivalent backbone size, how should the routing configuration be chosen? We develop a practical scaling-law framework by fitting X-MoD relative to FLOP-matched dense baselines, yielding an interpretable law that decomposes performance into sparse-capacity gain, sparse-context correction, and anchor-stride interaction. The law predicts validation loss across routing configurations and reveals how context length, model scale, and anchor stride shape sparse-depth performance. We validate the architecture and law through pretraining sweeps, held-out scaling-law prediction, ablations, downstream evaluations, and comparisons with Dense, MoD, and representative MoE baselines.
Figures & tables
Figure 1: MoD vs. X-MoD. X-MoD decouples token sparsity K from anchor stride A , enabling a substantially larger gap between total and active-equivalent parameter counts.
Figure 2: Empirical regularities motivating the X-MoD scaling law. (a, b) Loss is U-shaped in K , and the optimum shifts with context length L , also visible under sparse context Ls=L/K . (c) Varying Nact has weaker effect on optimal K . (d) Anchor stride A interacts with K . Sweep settings are given in Appendix B.4 .
Figure 3: Predicted vs. true validation loss. (a) In-range fit (circles) and out-of-range predictions without refitting (squares). (b–d) Leave-one-group-out predictions by L , A and Nact , respectively.
Nact
Model
N
ϕ/ϕD↓
Loss ↓
Hella. ↑
ARC-e ↑
ARC-c ↑
PIQA ↑
LAMB. ↑
BoolQ ↑
Avg. ↑
556M
Dense
0.54B
1.00
2.773
30.42
51.93
23.21
63.80
18.47
60.52
41.39
MoD
0.87B
0.92
2.692
34.00
58.75
25.94
66.43
21.33
51.99
43.07
MoE 8:64
2.46B
1.00
2.564
36.62
62.25
27.30
68.99
26.37
60.95
47.08
X-MoD A4K8
2.64B
0.47
2.549
38.30
62.08
27.56
68.34
29.19
58.62
47.35
MoE 8:128
4.58B
1.00
2.528
38.57
64.98
30.38
70.18
27.25
56.30
47.94
X-MoD A3K16
4.79B
0.46
2.503
38.00
64.31
28.41
70.13
30.49
60.49
48.64
Table 1: Main results under matched active-equivalent size and training compute. N denotes total parameters, and ϕ/ϕD denotes per-token training FLOPs normalized by the dense model with the same Nact and sequence length. The best results are highlighted in bold . Detailed compute budgets and full model configurations are reported in Appendix B .
Variant
Total N
ϕ/ϕD
Loss ↓
Hella. ↑
ARC-e ↑
ARC-c ↑
PIQA ↑
LAMB. ↑
BoolQ ↑
Avg. ↑
Dense
936M
1.00
2.635
35.18
59.85
26.71
67.52
24.16
59.85
45.54
MoD
1.48B
0.92
2.570
36.67
62.04
26.54
68.72
26.14
60.89
46.83
Routed-FFN A4K8
3.26B
1.00
2.508
37.97
64.27
28.92
70.29
29.79
60.95
48.70
Routed-FFN A3K16
5.68B
1.00
2.482
40.71
65.61
30.38
69.15
29.89
60.37
49.35
Full X-MoD A4K8
4.52B
0.51
2.427
41.26
68.69
33.45
71.22
33.55
60.43
51.43
Selected-Q / full-KV A4K8
4.52B
1.09
2.564
38.44
61.66
27.65
68.01
28.90
57.71
47.06
Table 2: Mechanism controls of X-MoD. All variants use the 936M reference backbone. Metrics follow Table 1 ; configurations are given in Appendix B.6 .
Variant
Loss ↓
Δ Loss
Hella. ↑
ARC-e ↑
ARC-c ↑
PIQA ↑
LAMB. ↑
BoolQ ↑
Avg. ↑
Dense reference
2.635
+0.159
35.18
59.85
26.71
67.52
24.16
59.85
45.54
Full X-MoD A1K16
2.476
0.000
40.64
66.08
30.80
71.27
30.58
61.31
50.12
w/o dense anchors
2.579
+0.103
36.35
62.63
28.33
69.53
26.59
58.96
47.06
w/o gated residual scaling
2.492
+0.016
40.62
65.19
29.27
69.31
29.73
60.58
49.12
w/o token-choice bias
2.657
+0.181
34.59
59.85
26.54
67.52
22.01
59.11
44.94
w/o dense prefix
2.487
+0.011
40.41
65.19
30.20
69.70
29.69
61.01
49.37
Table 3: Ablations of X-MoD stabilization mechanisms. All variants are compared under the same Nact=936 M. Δ Loss is measured relative to Full X-MoD; task scores follow Table 1 .
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Nact
Dim
Head dim
Layers
Heads
KV heads
FFN dim
165M
512
64
28
16
8
1536
298M
768
64
28
16
8
2304
556M
1024
128
28
16
8
3072
936M
1536
128
28
16
8
3840
1.65B
2048
128
28
16
8
6144
Appendix
Table 4: Backbone model configurations. All models use GQA attention and standard feed-forward layers.
Hyperparameter
Value
Peak learning rate
3.2×10−3
Momentum coefficient β
0.95
Newton–Schulz steps
5
Nesterov momentum
True
RMS match
0.2
AdamW betas
(0.9,0.95)
Appendix
Table 5: Muon optimizer configuration selected in preliminary dense-baseline sweeps and used in the main runs.
Nact
C
Model
Layers
Experts
de
N
FLOPs/tok.
Exec.
556M
1020
Dense
28
–
–
0.54B
8.4×109
DDP
MoD
49
–
–
0.87B
7.7×109
DDP
MoE 8:64
28
6:64+2
384
2.46B
8.4×109
DDP
X-MoD A4K8
161
–
–
2.64B
3.9×109
DDP
MoE 8:128
28
6:128+2
384
4.58B
8.4×109
DDP
X-MoD A3K16
298
–
–
4.79B
3.9×109
DDP
Appendix
Table 6: Detailed model configurations used in the main comparison. Nact denotes active-equivalent parameters, N denotes total parameters and C denotes total training FLOPs.
a, Routed computation and attention context
Model
Routed computation
KV context
Dense
None
Full
MoD
Attention + FFN
Selected subset
Routed-FFN A4K8
FFN only
Full every K refinements
Routed-FFN A3K16
FFN only
Full every K refinements
Selected-Q/full-KV A4K8
Q/O + FFN
Full sequence
Appendix
Table 7: Mechanism-control configurations at the 936M reference scale. D/DD=(ϕ/ϕD)−1 denotes token exposure relative to Dense.
Table 8: Fitted parameters of the practical X-MoD scaling law. For ease of reference, Eq. ( 14 ) is reproduced above the estimates.
Figure 4: Pilot comparison of X-MoD and MoE configurations. Validation loss versus total parameters at Nact=165M , L=32k and C=3×1019 FLOPs. Circles denote X-MoD and diamonds denote MoE; colour indicates K , and marker area increases with total parameters. Green dashed lines mark the selected MoE parameter budgets; black outlines identify the selected MoE–X-MoD pairs. The plot labels 8o64s2 and 8o128s2 correspond to MoE 8:64 and 8:128.
Variant
Total N
N/Nact
Total layers
ϕ/ϕD
Val. loss ↓
Deep Dense
256M
1.00
52
1.00
3.084
Deep X-MoD A1K4
552M
2.16
124
0.67
3.047
Deep X-MoD A1K8
939M
3.67
220
0.62
3.034
Deep X-MoD A1K12
1.29B
5.04
316
0.60
3.023
Deep X-MoD A1K16
1.67B
6.52
412
0.59
3.011
Appendix
Table 9: Total-depth scaling of X-MoD. All variants use the width of the 165M backbone, with matched active-equivalent capacity and training compute C=4×1019 .
Repeated identical subset
Independent random subsets
Threshold routing
Threshold / top- k
Model
Coverage
Jaccard
Coverage
Jaccard
Coverage
Jaccard
Agreement
X-MoD A1K16
6.250
100.0
64.39
3.226
52.24
14.51
98.63
A1K16 w/o token-choice bias
6.250
100.0
64.39
3.226
58.59
80.66
–
X-MoD A4K8
12.50
100.0
98.61
6.667
99.99
1.947
96.62
X-MoD A3K16
6.250
100.0
95.49
3.226
97.50
2.940
98.25
Appendix
Table 10: Routing coverage, consecutive-layer overlap and mask agreement. All entries are percentages at the 936M reference scale; a dash denotes an unmeasured value.
Figure 5: Token-wise sparse-update distributions under threshold routing. (a) A1K16 with and without token-choice bias (16-layer intervals). (b) A4K8 (complete 32-layer intervals). (c) A3K16 (48-layer intervals). Grey dashed and green dot-dashed curves show repeated-subset and independent-subset references. Curves interpolate discrete probabilities; integer heights, not filled areas, represent probability. Models and evaluation conditions follow Table 10 .
a, Model scale, counted compute and peak training memory
Model
Layers
N
FLOPs/token ↓
Peak memory (GB/GPU) ↓
Dense
28
936M
9.0G
32.08
MoD
49
1.48B
8.3G
46.62
MoE 8:64
28
4.51B
9.0G
55.46
X-MoD A4K8
161
4.52B
4.6G
55.72
MoE 8:128
28
8.48B
9.0G
71.96
Appendix
Table 11: Measured systems efficiency at the 936M active-equivalent scale. Speedups are relative to the adjacent, approximately total-parameter-matched MoE baseline.