Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In this work, we introduce Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its core is a bounded, sparsity-conditional recurrence mapping that characterizes the effective-parameter gain from looping and how sparsity raises this gain. The laws predict the held-out loss of looped models more accurately than prior alternatives, and recover the standard dense and MoE scaling laws as special cases. Beyond prediction, the fitted laws provide a principled foundation for designing looped MoE models under compute and memory constraints. Downstream evaluations further demonstrate the complementary benefits of the two axes: sparsity delivers ~3x active-parameter efficiency, recurrence yields ~2x total-parameter efficiency on reasoning, and joint scaling further advances the performance frontier. As a practical extension, we show these gains hold at trillion-token scale: at matched training compute, a looped MoE with law-derived recurrence matches a ~2x larger non-looped MoE on the reasoning benchmarks, while enabling test-time scaling through recurrence.
Figures & tables
Figure 1: (a) Dense looped transformers reuse the same FFNs, whereas looped MoE models can route tokens to different experts across recurrent passes. (b) At fixed active parameters and training tokens, increasing recurrence yields diminishing returns, with the dense model ( E=1 ) saturating earlier than MoE models ( E>1 ). (c) Our Loop Scaling Laws model a bounded effective-parameter gain conditioned on recurrence R and sparsity (varied by expert count E ), unlike prior linear and power-law mappings that assume unbounded gains.
Scaling axes
Optimality
LLM Category
Scaling law
N/F
D
R
S/E
Recurrence mapping
Compute
Memory
Dense
Chinchilla ( 2022 )
✓
✓
×
×
–
✓
×
MoE
Unified routed ( 2022 )
✓
×
×
✓
–
×
×
Joint MoE ( 2025 )
✓
✓
×
✓
–
✓
✓
Looped
Parcae ( 2026 /04 )
✓
✓
✓
×
linear; unbounded †
✓
×
Iso-Depth ( 2026 /05 )
✓
✓
✓
×
power law; unbounded
✓
×
Table 1: Comparison of scaling laws for training LLMs. ✓ indicates variables explicitly modeled by each law. N/F denotes model size N / training FLOPs F . S/E denotes MoE sparsity S (varied by expert expansion E ).
Figure 2: (a) Illustration of effective-parameter gain under linear, power-law, bounded recurrence mappings. (b) Predictive loss by different recurrence mappings: each mapping is fitted on R≤8 and applied to predict the held-out- R up to R=16 . (c) Formula on recurrence mappings and their effective-parameter gains (top) and their evaluation on held-out recurrence R , model N , and data D with RMSE scores (bottom).
Figure 3: (a) Illustration of effective-parameter gain with sparsity-conditional recurrence mapping. (b) Predictive loss under different MoE recurrence mappings fitted on runs with R≤8 ; markers denote observations used for fitting. (c) Formula on MoE recurrence mappings and their effective-parameter gains (top), and their evaluation on held-out recurrence R , expert count E , model size N , and data D with RMSE scores (bottom).
Figure 4: (a,b) Effective-parameter multiplier over recurrence at Nact=1.0 B, and its asymptote over expert count. (c,d) Recurrence IsoFLOP profiles for dense ( E=1 ) and MoE ( E=8 ) models; markers denote minima.
Figure 5: Predicted compute-optimal recurrence R⋆ of dense ( E=1 ) and MoE ( E=8 ) models. (a) Recurrence with compute-optimal loss across training-compute budgets. (b) IsoFLOP profiles over recurrence at Fˉtrain=5×1021 FLOPs; stars mark R⋆ . (c) R⋆ across compute and active model size for E=8 (top) and E=1 (bottom).
Figure 6: Predicted memory-optimal expert count E⋆ and joint optimum (Nact⋆,E⋆,R⋆) at Fˉtrain=5×1021 FLOPs. (a) E⋆ across bf16 weight-memory budgets at R=4 (top) and R=1 (bottom). (b) E⋆ across INT4 weight memory and recurrence. (c) Predicted joint optimum (Nact⋆,E⋆,R⋆) across INT4 weight memory.
Figure 7: Downstream scaling of sparsity and recurrence. (a,b) Scaling E at R=1 ; (c,d) Scaling R at E=8 , with the dense baseline E=1,R=1 . (a,c) report Overall performance, and (b,d) report Reasoning performance; dashed lines mark compute-matched comparisons.
Model
R
Finf0.6B(1)Finf0.3B(R)
Reasoning
Science
Commonsense Reasoning
Reading
Knowledge
BBH 3
GSM8K 8
ARC-C 25
ARC-E
OBQA
HS
PIQA
SIQA
Wino
BoolQ
DROP 3
MMLU 5
NQ 5
TQA 5
Overall ( Δ )
A0.6B-2.9B MoE
1
1.0×
29.8
42.9
50.0
66.9
40.8
68.1
77.2
51.6
64.4
70.1
39.7
46.9
13.8
39.8
50.2
A0.3B-1.3B LoopMoE
1
0.4×
18.6
10.5
36.1
57.3
33.8
52.7
70.7
43.8
54.6
58.1
23.0
29.2
5.7
18.2
36.6
2
0.8×
26.7
30.0
42.2
60.3
38.2
59.9
72.5
46.4
58.5
62.8
33.7
37.6
8.3
24.1
43.0 (+6.4)
3
1.1×
30.1
41.0
44.9
60.9
39.6
62.4
73.2
50.2
61.3
65.9
38.4
41.6
8.9
25.6
46.0 (+9.4)
4
1.5×
31.5
41.5
46.5
63.9
39.0
62.9
73.9
51.1
60.6
64.4
39.3
42.2
9.8
27.1
46.7 (+10.1)
Table 2: Downstream results of A0.6B-2.9B MoE and A0.3B-1.3B LoopMoE across test-time recurrence R , trained at matched compute FLOPs ≈1.5×1022 . Relative inference compute is normalized to the non-looped baseline. Δ reports gains over LoopMoE at R=1 . Evaluation protocols are provided in Appendix D.1 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
E
R=1
R=2
R=4
R=8
1
1.0000±0.0000
1.0000±0.0000
1.0000±0.0000
1.0000±0.0000
2
1.0000±0.0000
1.1212±0.0692
1.2165±0.0465
1.5088±0.0529
4
1.0000±0.0000
1.1917±0.0974
1.2973±0.0703
1.5300±0.0785
8
1.0000±0.0000
1.2099±0.0868
1.4159±0.0901
1.7035±0.0861
Appendix
Table B.1: Expert-path diversity Ψ across varying recurrence R and expert count E (mean ± std).
Architecture dimensions
Ntotal(E) (B)
Nact (B)
dm
nh
nℓ
nℓ,loop
Nloop (B)
E=1
E=2
E=4
E=8
E=16
0.3
768
12
20
16
0.1
0.3
0.5
0.8
1.3
2.5
0.6
1024
16
26
22
0.3
0.6
0.9
1.6
2.9
5.5
1.0
1280
20
32
28
0.7
1.0
1.6
2.9
5.4
10.5
1.6
1536
24
38
34
1.2
1.6
2.7
4.8
9.1
17.7
2.4
1792
28
44
40
1.8
2.4
4.1
7.5
14.3
27.8
Appendix
Table C.1: Model scaling ladder over five model scales. Parameter counts are in billions (B). Nact and Ntotal include embedding parameters, whereas Nloop includes only parameters in the recurrent block.
Base scaling law coefficients
MoE coefficients
Recurrence coefficients
Fit
A
α
B
β
c
δ
γ
ω
ζ
Estart
Emax
κ1
κ2
θ
RMSE
0.8085
−0.1951
3.4291
−0.7548
1.3555
−0.1935
−0.0137
−0.2056
0.0881
1.3174
57.1201
0.3682
1.4037
0.3286
0.0037
Appendix
Table C.2: Fitted coefficients of MoE Loop Scaling Law in Eq. 9 . RMSE: root-mean-square error of the fit.
Coefficient
Estimate
90% CI
κ1
0.3682
[0.3658,0.3723]
κ2
1.4037
[1.3896,1.4243]
θ
0.3286
[0.3183,0.3350]
Appendix
Table C.3: Bootstrap results on the MoE loop scaling law. (a) Coefficients fitted on full-data or the 90% sampled subsets with intervals spanning the 5th–95th percentiles over 100 refit runs. (b) Derived joint optimum under the compute and memory budgets (with INT4 weights) used in Figure 6 (c) on the 100 refit runs.
Category
Benchmark
Task name
n -shot
Metric
Reasoning
BBH
leaderboard_bbh
3
acc_norm
GSM8K
gsm8k_cot
8
exact_match,flexible-extract
Science
ARC-C
arc_challenge
25
acc_norm
ARC-E
arc_easy
0
acc_norm
OBQA
openbookqa
0
acc_norm
Commonsense Reasoning
HellaSwag
hellaswag
0
acc_norm
Appendix
Table D.1: Evaluation setups of 14 foundational benchmarks. Task names and metrics follow lm-eval .