cs.LGOct 1, 2026

Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

Authors: Di He, Pengxiang Li, Da Chang, Qingyan Meng, Lu Yin, Shiwei Liu

Organizations: Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences · Peng Cheng Laboratory · University of Chinese Academy of Sciences · The Hong Kong Polytechnic University · University of Surrey · ELLIS Institute Tübingen · Max Planck Institute for Intelligent Systems · Tübingen AI Center

Abstract

Looped Transformers introduce recurrent depth as a new scaling axis for LLMs: by repeatedly applying shared Transformer blocks, they increase effective depth without increasing parameter count. However, the benefits of looping remain unclear for large MoE LLMs under FLOPs-matched comparisons. The main reason is that the gains from additional iterations diminish quickly and can even turn into degradation, so the extra FLOPs spent on looping yield little substantial improvement. Consequently, prior work typically settles on two loops. We identify two main obstacles to scaling looped MoE. First, looping inherits and amplifies the curse of depth: hidden-state variance grows with each iteration as residual updates accumulate, which destabilizes deep recurrence and causes representations to drift. Second, looped MoE suffers from expert selection collapse: routers repeatedly select the same experts across loops, so extra iterations add computation without adding computational diversity. Guided by this diagnosis, we propose LOOM, built on a single principle: each loop should contribute new computation while keeping the recurrent state stable. LOOM stabilizes recurrence by scaling residual updates to bound variance growth and re-injecting the input embedding at every loop, and diversifies it through per-loop routers that engage different experts and a Looping Residual that carries earlier outputs forward. Experiments across 100M-1.7B models show stable scaling to 9-12 loops. Under near-iso-FLOP, the 700M model performs best at 5 loops, reducing perplexity from 18.36 to 16.54 and improving average zero-shot accuracy from 38.84% to 39.53% over the non-looped baseline. Without FLOP matching, the 1.7B model trained on 60B tokens peaks at 9 loops, reducing perplexity from 9.62 to 7.77 and improving average zero-shot accuracy from 42.4% to 47.7%. Code is available https://github.com/hed-ucas/LOOM.

Figures & tables

Explore similar work

CardsList
  1. Sparse Layers are Critical to Scaling Looped Language Models

    May 9, 2026Ryan Lee, Jacob Biloki, Edward J. Hu +1Transformer ArchitecturesLayer-Wise

  2. How to Loop MoE: Flatten the Experts, Untie the Attention

    Sep 28, 2026Shouren Wang, Chuang Ma, Mohsen Hariri +6Mixture-Of-ExpertsExperts

  3. SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

    Date pendingShaowen Wang, Ge Zhang, Kairong Luo +6Transformer ArchitecturesScaling Laws