Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts
Authors: Di He, Pengxiang Li, Da Chang, Qingyan Meng, Lu Yin, Shiwei Liu
Organizations: Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences · Peng Cheng Laboratory · University of Chinese Academy of Sciences · The Hong Kong Polytechnic University · University of Surrey · ELLIS Institute Tübingen · Max Planck Institute for Intelligent Systems · Tübingen AI Center
Looped Transformers introduce recurrent depth as a new scaling axis for LLMs: by repeatedly applying shared Transformer blocks, they increase effective depth without increasing parameter count. However, the benefits of looping remain unclear for large MoE LLMs under FLOPs-matched comparisons. The main reason is that the gains from additional iterations diminish quickly and can even turn into degradation, so the extra FLOPs spent on looping yield little substantial improvement. Consequently, prior work typically settles on two loops. We identify two main obstacles to scaling looped MoE. First, looping inherits and amplifies the curse of depth: hidden-state variance grows with each iteration as residual updates accumulate, which destabilizes deep recurrence and causes representations to drift. Second, looped MoE suffers from expert selection collapse: routers repeatedly select the same experts across loops, so extra iterations add computation without adding computational diversity. Guided by this diagnosis, we propose LOOM, built on a single principle: each loop should contribute new computation while keeping the recurrent state stable. LOOM stabilizes recurrence by scaling residual updates to bound variance growth and re-injecting the input embedding at every loop, and diversifies it through per-loop routers that engage different experts and a Looping Residual that carries earlier outputs forward. Experiments across 100M-1.7B models show stable scaling to 9-12 loops. Under near-iso-FLOP, the 700M model performs best at 5 loops, reducing perplexity from 18.36 to 16.54 and improving average zero-shot accuracy from 38.84% to 39.53% over the non-looped baseline. Without FLOP matching, the 1.7B model trained on 60B tokens peaks at 9 loops, reducing perplexity from 9.62 to 7.77 and improving average zero-shot accuracy from 42.4% to 47.7%. Code is available https://github.com/hed-ucas/LOOM.
Figures & tables
Figure 1: Loop-depth scaling of LOOM . Left: under a near-iso-FLOP budget, LOOM achieves its best performance near 5 loops. Right: without FLOP matching, the 1.7B model scales to 9 loops, with substantial performance improvements.
Figure 2: LOOM architecture. An M -layer MoE block is reused for H loops, with λ as a scalar hyperparameter. At layer a of loop t , γ=λ/(HM) scales residual updates, gt=λ/(tM) controls embedding re-injection, EMAat denotes the Looping Residual, and Rat denotes the router.
Figure 3: Effect of residual scaling and embedding re-injection on training stability in a 9-loop MoE model ( ∼350 M parameters, 10 B training tokens). Native (No Tech) loops the unmodified backbone without additional techniques. Only Res Scale and Only Embed Inject add residual scaling and embedding re-injection, respectively, to Native, with no other modifications. (a) Training loss. (b) Activation variance during training. (c) Activation variance across recurrent iterations.
Figure 4: Expert routing in a 9-loop MoE model ( ∼350 M parameters), comparing loop-specific (Independent) and shared (Shared) routers. (a) Cross-loop similarity: cos(pˉt,pˉt′) , where pˉt is the normalized expert-load distribution averaged across layers for loop t . Each cell compares two loops; higher values indicate more similar expert utilization. (b) Cross-loop variability: stdt(pℓ,t,e) , where pℓ,t,e is the normalized load of expert e at layer ℓ in loop t . Each cell measures an expert’s load variation across loops at a given layer. Loop-specific routers yield lower similarity and greater variability, indicating more diverse expert utilization.
Figure 5: Effect of segmented backpropagation ( K=3 ) on the ∼350 M model, compared with backpropagation through all loops. (a) Peak NPU memory measured with a microbatch size of 1. (b) Total training times. (c) Held-out evaluation loss (solid bars) and average accuracy across seven downstream tasks (hatched bars).
Parameters (activated)
∼100 M ( ∼70 M)
∼350 M ( ∼141 M)
∼1.7 B ( ∼0.63 B)
Physical layers M
6
10
15
Effective depth
6H
10H
15H
Hidden size d
384
512
1280
Query / KV heads
6 / 3
8 / 4
20 / 10
Head dimension
64
64
64
Expert width I
384
512
768
Table 1: Architecture and optimization settings for the three LOOM scales. Effective depth is MH , with an M -layer block reused for H recurrent iterations.
Method
Loops
Eff. depth
top- k
f(H,k)
Val.ppl
OBQA
Wino
ARCC
ARCE
HSwg
SIQA
PIQA
Avg.
Baseline
1×
10
26
84
18.36
28.40
51.38
22.95
43.10
29.43
35.72
60.88
38.84
LOOM
2×
20
12
84
17.37
30.60
49.33
24.32
44.57
29.98
34.65
59.58
39.00
LOOM
3×
30
8
90
16.91
30.20
50.38
25.00
44.02
30.05
35.18
60.55
39.34
LOOM
4×
40
5
84
16.58
30.20
50.28
24.74
44.11
30.14
36.01
61.26
39.53
LOOM
5×
50
4
90
16.54
30.80
50.20
24.91
44.70
30.46
34.65
60.99
39.53
LOOM
6×
60
3
90
16.57
30.40
52.17
24.66
44.02
29.91
35.16
60.17
39.50
Table 2: Near-iso-FLOP comparison on the M=10 , E=80 backbone. All models have 700M parameters and are trained on 10B tokens. Per-token layer cost is f=H(6+3k) in units of d2 , anchored at f1=84 . Best results are bold.
Model size
Method
Loops
Eff. depth
top- k
Val.ppl
OBQA
Wino
ARCC
ARCE
HSwg
SIQA
PIQA
Avg.
Baseline
1×
6
6
26.26
27.20
48.46
22.95
40.07
26.85
34.54
57.02
36.73
No tech.
3×
18
6
27.61
27.80
48.30
21.93
38.89
26.24
34.03
56.96
36.31
No tech.
6×
36
6
1311†
27.40
51.54
24.49
27.53
25.51
34.90
48.91
34.33†
No tech.
9×
54
6
1312†
26.80
47.75
23.63
27.27
25.61
34.90
48.97
33.56†
No tech.
12×
72
6
1587†
27.20
49.33
24.06
28.07
25.34
34.18
49.51
33.96†
Embed inject
3×
18
6
27.48
24.80
52.01
22.27
37.75
26.21
33.73
54.68
35.92
Table 3: Non-iso-FLOP loop-depth sweeps at ∼100 M and ∼350 M scales, including partial stabilization baselines. We report validation perplexity ( ↓ ) and zero-shot accuracy (%, ↑ ). Best results within each scale are bold; † denotes an unrecoverable loss spike.
Method
Loops
Eff. depth
top- k
Val.ppl
OBQA
Wino
ARCC
ARCE
HSwg
SIQA
PIQA
Avg.
Baseline
1×
15
8
9.62
29.6
50.3
26.7
50.2
37.1
37.2
66.0
42.4
LOOM
3×
45
8
8.94
31.6
51.3
27.2
51.0
39.1
39.1
68.0
43.9
LOOM
6×
90
8
7.91
32.2
51.5
31.1
57.0
46.6
39.5
69.2
46.7
LOOM
9×
135
8
7.77
33.0
53.6
31.7
57.7
48.8
39.2
70.2
47.7
LOOM
12×
180
8
7.84
34.8
50.7
30.5
56.9
48.6
40.0
70.7
47.5
Table 4: Scaling LOOM to ∼1.7 B parameters and ∼60 B tokens. We report validation perplexity ( ↓ ) and zero-shot accuracy (%, ↑ ).
Variant
Val.ppl
OBQA
Wino
ARCC
ARCE
HSwg
SIQA
PIQA
Avg.
LOOM
18.62
31.4
51.1
26.1
43.3
29.5
35.8
61.4
39.8
w/o Res. scale
24.63
29.2
50.1
23.8
42.6
29.4
35.9
59.4
38.6
w/o Embed inject
26.05
29.4
50.2
24.3
43.9
29.2
36.2
61.5
39.2
w/o MoE RMSNorm
26.50
29.6
52.1
24.1
43.6
28.7
35.5
59.4
39.0
w/o Looping Residual
19.25
30.8
49.2
24.2
44.8
29.5
36.6
60.2
39.3
w/o routing refresh
19.83
30.0
51.7
25.4
43.1
29.4
35.7
59.2
39.2
Table 5: Component ablation of the ∼350 M LOOM model with 9 loops, evaluated at training step 5,000. Each ablated variant removes one component from the full recipe; “w/o routing refresh” uses a shared router across loops, and “w/o LR” removes the Looping Residual. The best result in each column is shown in bold.
Loops
Backprop
Val.ppl
OBQA
Wino
ARCC
ARCE
HSwg
SIQA
PIQA
Avg.
K=3
18.27
31.2
51.7
24.2
44.2
29.0
35.1
60.3
39.4
3×
none
26.63
29.2
50.2
23.1
43.7
28.4
35.2
60.8
38.7
K=3
15.35
32.2
50.4
26.1
48.0
31.6
36.1
61.7
40.9
6×
none
25.33
29.2
51.5
24.7
43.2
29.0
35.7
60.3
39.1
K=3
14.80
32.0
50.1
27.2
47.2
32.4
36.9
62.2
41.1
9×
none
81.04†
25.2
47.5
20.6
32.9
25.1
35.1
53.4
34.3†
Table 6: Comparison of segmented backpropagation ( K=3 ) and full backpropagation across loop depths on the ∼350 M model.