cs.LGSep 28, 2026

How to Loop MoE: Flatten the Experts, Untie the Attention

Authors: Shouren Wang, Chuang Ma, Mohsen Hariri, Debargha Ganguly, Wang Yang, Xiaoqing Tong, Qianying Liu, Xiaotian Han, +1 more

Organizations: Case Western Reserve University · Kyoto University · NII LLMC

Abstract

Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each token. Looped MoE bridges these two design philosophies and gives MoE models new potential for better expert usage, but it raises a question: how to loop a MoE? We answer it with Foil. With the expert parameters and the expert compute per token held fixed, Foil (1) flattens the experts, halving the expert layers, doubling the experts per layer and doubling the passes, so that every routing decision chooses from a larger pool, and (2) unties the attention, giving each pass its own attention parameters while the experts and routers stay shared. Experiments show that Foil clearly outperforms the unflattened looped baseline: at 20B tokens every Foil model has lower pretraining loss than the baseline; at 100B tokens the loss improves monotonically with the degree of flattening, the most flattened Foil ending 0.012 nat below the baseline at equal parameters and compute, with downstream accuracy on par or better; untying the attention also yields more balanced and more confident routing at equal shape. Our ablations analyse why Foil works and turn the findings into design guidance for looped MoE: the returns of looping and of widening the expert layers amplify each other, routing confidence tracks healthy expert use better than load balance, and a sparse looped MoE should therefore use more experts per layer and more passes. Code and configurations are available at https://github.com/SR-A-W/how-to-loop-moe.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

    Oct 1, 2026Di He, Pengxiang Li, Da Chang +3Mixture-Of-Experts Large Language ModelsMixture-Of-Experts

  2. Sparse Layers are Critical to Scaling Looped Language Models

    May 9, 2026Ryan Lee, Jacob Biloki, Edward J. Hu +1Transformer ArchitecturesLayer-Wise

  3. SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

    Date pendingShaowen Wang, Ge Zhang, Kairong Luo +6Transformer ArchitecturesScaling Laws