Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts
Authors: Di He, Pengxiang Li, Da Chang, Qingyan Meng, Lu Yin, Shiwei Liu
Organizations: Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences · Peng Cheng Laboratory · University of Chinese Academy of Sciences · The Hong Kong Polytechnic University · University of Surrey · ELLIS Institute Tübingen · Max Planck Institute for Intelligent Systems · Tübingen AI Center
Looped Transformers introduce recurrent depth as a new scaling axis for LLMs: by repeatedly applying shared Transformer blocks, they increase effective depth without increasing parameter count. However, the benefits of looping remain unclear for large MoE LLMs under FLOPs-matched comparisons. The main reason is that the gains from additional iterations diminish quickly and can even turn into degradation, so the extra FLOPs spent on looping yield little substantial improvement. Consequently, prior work typically settles on two loops. We identify two main obstacles to scaling looped MoE. First, looping inherits and amplifies the curse of depth: hidden-state variance grows with each iteration as residual updates accumulate, which destabilizes deep recurrence and causes representations to drift. Second, looped MoE suffers from expert selection collapse: routers repeatedly select the same experts across loops, so extra iterations add computation without adding computational diversity. Guided by this diagnosis, we propose LOOM, built on a single principle: each loop should contribute new computation while keeping the recurrent state stable. LOOM stabilizes recurrence by scaling residual updates to bound variance growth and re-injecting the input embedding at every loop, and diversifies it through per-loop routers that engage different experts and a Looping Residual that carries earlier outputs forward. Experiments across 100M-1.7B models show stable scaling to 9-12 loops. Under near-iso-FLOP, the 700M model performs best at 5 loops, reducing perplexity from 18.36 to 16.54 and improving average zero-shot accuracy from 38.84% to 39.53% over the non-looped baseline. Without FLOP matching, the 1.7B model trained on 60B tokens peaks at 9 loops, reducing perplexity from 9.62 to 7.77 and improving average zero-shot accuracy from 42.4% to 47.7%. Code is available https://github.com/hed-ucas/LOOM.
Figures & tables
Figure 1: Loop-depth scaling of LOOM . Left: under a near-iso-FLOP budget, LOOM achieves its best performance near 5 loops. Right: without FLOP matching, the 1.7B model scales to 9 loops, with substantial performance improvements.
Figure 2: LOOM architecture. An M -layer MoE block is reused for H loops, with λ as a scalar hyperparameter. At layer a of loop t , γ=λ/(HM) scales residual updates, gt=λ/(tM) controls embedding re-injection, EMAat denotes the Looping Residual, and Rat denotes the router.
Figure 3: Effect of residual scaling and embedding re-injection on training stability in a 9-loop MoE model ( ∼350 M parameters, 10 B training tokens). Native (No Tech) loops the unmodified backbone without additional techniques. Only Res Scale and Only Embed Inject add residual scaling and embedding re-injection, respectively, to Native, with no other modifications. (a) Training loss. (b) Activation variance during training. (c) Activation variance across recurrent iterations.
Figure 4: Expert routing in a 9-loop MoE model ( ∼350 M parameters), comparing loop-specific (Independent) and shared (Shared) routers. (a) Cross-loop similarity: cos(pˉt,pˉt′) , where pˉt is the normalized expert-load distribution averaged across layers for loop t . Each cell compares two loops; higher values indicate more similar expert utilization. (b) Cross-loop variability: stdt(pℓ,t,e) , where pℓ,t,e is the normalized load of expert e at layer ℓ in loop t . Each cell measures an expert’s load variation across loops at a given layer. Loop-specific routers yield lower similarity and greater variability, indicating more diverse expert utilization.
Figure 5: Effect of segmented backpropagation ( K=3 ) on the ∼350 M model, compared with backpropagation through all loops. (a) Peak NPU memory measured with a microbatch size of 1. (b) Total training times. (c) Held-out evaluation loss (solid bars) and average accuracy across seven downstream tasks (hatched bars).
Parameters (activated)
∼100 M ( ∼70 M)
∼350 M ( ∼141 M)
∼1.7 B ( ∼0.63 B)
Physical layers M
6
10
15
Effective depth
6H
10H
15H
Hidden size d
384
512
1280
Query / KV heads
6 / 3
8 / 4
20 / 10
Head dimension
64
64
64
Expert width I
384
512
768
Table 1: Architecture and optimization settings for the three LOOM scales. Effective depth is MH , with an M -layer block reused for H recurrent iterations.
Method
Loops
Eff. depth
top- k
f(H,k)
Val.ppl
OBQA
Wino
ARCC
ARCE
HSwg
SIQA
PIQA
Avg.
Baseline
1×
10
26
84
18.36
28.40
51.38
22.95
43.10
29.43
35.72
60.88
38.84
LOOM
2×
20
12
84
17.37
30.60
49.33
24.32
44.57
29.98
34.65
59.58
39.00
LOOM
3×
30
8
90
16.91
30.20
50.38
25.00
44.02
30.05
35.18
60.55
39.34
LOOM
4×
40
5
84
16.58
30.20
50.28
24.74
44.11
30.14
36.01
61.26
39.53
LOOM
5×
50
4
90
16.54
30.80
50.20
24.91
44.70
30.46
34.65
60.99
39.53
LOOM
6×
60
3
90
16.57
30.40
52.17
24.66
44.02
29.91
35.16
60.17
39.50
Table 2: Near-iso-FLOP comparison on the M=10 , E=80 backbone. All models have 700M parameters and are trained on 10B tokens. Per-token layer cost is f=H(6+3k) in units of d2 , anchored at f1=84 . Best results are bold.
Model size
Method
Loops
Eff. depth
top- k
Val.ppl
OBQA
Wino
ARCC
ARCE
HSwg
SIQA
PIQA
Avg.
Baseline
1×
6
6
26.26
27.20
48.46
22.95
40.07
26.85
34.54
57.02
36.73
No tech.
3×
18
6
27.61
27.80
48.30
21.93
38.89
26.24
34.03
56.96
36.31
No tech.
6×
36
6
1311†
27.40
51.54
24.49
27.53
25.51
34.90
48.91
34.33†
No tech.
9×
54
6
1312†
26.80
47.75
23.63
27.27
25.61
34.90
48.97
33.56†
No tech.
12×
72
6
1587†
27.20
49.33
24.06
28.07
25.34
34.18
49.51
33.96†
Embed inject
3×
18
6
27.48
24.80
52.01
22.27
37.75
26.21
33.73
54.68
35.92
Table 3: Non-iso-FLOP loop-depth sweeps at ∼100 M and ∼350 M scales, including partial stabilization baselines. We report validation perplexity ( ↓ ) and zero-shot accuracy (%, ↑ ). Best results within each scale are bold; † denotes an unrecoverable loss spike.
Method
Loops
Eff. depth
top- k
Val.ppl
OBQA
Wino
ARCC
ARCE
HSwg
SIQA
PIQA
Avg.
Baseline
1×
15
8
9.62
29.6
50.3
26.7
50.2
37.1
37.2
66.0
42.4
LOOM
3×
45
8
8.94
31.6
51.3
27.2
51.0
39.1
39.1
68.0
43.9
LOOM
6×
90
8
7.91
32.2
51.5
31.1
57.0
46.6
39.5
69.2
46.7
LOOM
9×
135
8
7.77
33.0
53.6
31.7
57.7
48.8
39.2
70.2
47.7
LOOM
12×
180
8
7.84
34.8
50.7
30.5
56.9
48.6
40.0
70.7
47.5
Table 4: Scaling LOOM to ∼1.7 B parameters and ∼60 B tokens. We report validation perplexity ( ↓ ) and zero-shot accuracy (%, ↑ ).
Variant
Val.ppl
OBQA
Wino
ARCC
ARCE
HSwg
SIQA
PIQA
Avg.
LOOM
18.62
31.4
51.1
26.1
43.3
29.5
35.8
61.4
39.8
w/o Res. scale
24.63
29.2
50.1
23.8
42.6
29.4
35.9
59.4
38.6
w/o Embed inject
26.05
29.4
50.2
24.3
43.9
29.2
36.2
61.5
39.2
w/o MoE RMSNorm
26.50
29.6
52.1
24.1
43.6
28.7
35.5
59.4
39.0
w/o Looping Residual
19.25
30.8
49.2
24.2
44.8
29.5
36.6
60.2
39.3
w/o routing refresh
19.83
30.0
51.7
25.4
43.1
29.4
35.7
59.2
39.2
Table 5: Component ablation of the ∼350 M LOOM model with 9 loops, evaluated at training step 5,000. Each ablated variant removes one component from the full recipe; “w/o routing refresh” uses a shared router across loops, and “w/o LR” removes the Looping Residual. The best result in each column is shown in bold.
Loops
Backprop
Val.ppl
OBQA
Wino
ARCC
ARCE
HSwg
SIQA
PIQA
Avg.
K=3
18.27
31.2
51.7
24.2
44.2
29.0
35.1
60.3
39.4
3×
none
26.63
29.2
50.2
23.1
43.7
28.4
35.2
60.8
38.7
K=3
15.35
32.2
50.4
26.1
48.0
31.6
36.1
61.7
40.9
6×
none
25.33
29.2
51.5
24.7
43.2
29.0
35.7
60.3
39.1
K=3
14.80
32.0
50.1
27.2
47.2
32.4
36.9
62.2
41.1
9×
none
81.04†
25.2
47.5
20.6
32.9
25.1
35.1
53.4
34.3†
Table 6: Comparison of segmented backpropagation ( K=3 ) and full backpropagation across loop depths on the ∼350 M model.
Looped language models repeat a set of transformer layers through depth, reducing memory costs and providing natural early-exit points at loop boundaries. However, looped models do not scale as favorably as standard transformers with unique layers. We compare standard and Mixture-of-Experts (MoE) transformers, with and without looping, and find two main results. First, we find Looped-MoE models scale better than the standard baseline while dense looped models do not. We trace this to routing divergence between loops: in Looped-MoE models, different experts are activated on each pass through the same shared layers, recovering expressivity without additional parameters. Our second finding is that looped models have better compute-quality trade-offs with early exits than standard models. Because each loop ends with the same layers that produce the final output, loop boundaries are superior exit points, as confirmed by earlier output convergence at these points. In sum, we provide a clear direction for scaling looped models: a Looped-MoE model with early exits can not only beat standard transformers at scale, but also enable significant memory and inference savings with minimal degradation in quality.
Ryan Lee, Jacob Biloki, Edward J. Hu +1
1USC Information Sciences Institute · 2Netflix · 3Independent Researcher
Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each token. Looped MoE bridges these two design philosophies and gives MoE models new potential for better expert usage, but it raises a question: how to loop a MoE? We answer it with Foil. With the expert parameters and the expert compute per token held fixed, Foil (1) flattens the experts, halving the expert layers, doubling the experts per layer and doubling the passes, so that every routing decision chooses from a larger pool, and (2) unties the attention, giving each pass its own attention parameters while the experts and routers stay shared. Experiments show that Foil clearly outperforms the unflattened looped baseline: at 20B tokens every Foil model has lower pretraining loss than the baseline; at 100B tokens the loss improves monotonically with the degree of flattening, the most flattened Foil ending 0.012 nat below the baseline at equal parameters and compute, with downstream accuracy on par or better; untying the attention also yields more balanced and more confident routing at equal shape. Our ablations analyse why Foil works and turn the findings into design guidance for looped MoE: the returns of looping and of widening the expert layers amplify each other, routing confidence tracks healthy expert use better than load balance, and a sparse looped MoE should therefore use more experts per layer and more passes. Code and configurations are available at https://github.com/SR-A-W/how-to-loop-moe.
Shouren Wang, Chuang Ma, Mohsen Hariri +6
Case Western Reserve University · Kyoto University · NII LLMC
Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating architectural advantage with extra FLOPs. We study looping on Mixture-of-Experts Transformers while closely matching per-token FLOPs, total non-embedding parameters, and KV cache. Through a series of ablations, we arrive at a recipe we call SMELT (Sparse MoE Transformer, middle layers Loop Twice), which loops the middle half of layers twice while matching the unlooped Baseline on all three budgets. We scale SMELT across four sizes up to 54B non-embedding parameters and fit a separate Chinchilla-style scaling law for each architecture. SMELT's loss drops faster with compute, saving 6.8--18.0% of training FLOPs on the compute-optimal frontier. The advantage transfers to downstream benchmarks beyond what validation loss predicts, is largest on Code, and grows with sample length and the number of in-context examples. Mechanistic analysis shows that the second visit reduces the attention sink and redirects mass toward content-relevant tokens, an inductive bias that may underlie the observed performance gains. These results show that looping can improve Transformers even under budget matching, offering a practical recipe that turns depth reuse into measurable gains.