cs.LGSep 28, 2026

Loop Dropout: Regularizing Shared Updates in Looped Language Models

Authors: Zirui Zhu, Hailun Xu, Xuanlei Zhao, Yong Liu, Yingxuan Ren, Kanchan Sarkar, Kun Xu, Yang You

Abstract

Looped language models separate computational depth from parameter count by repeatedly applying the same transformer block. Adapting these models requires a shared update that remains effective as hidden states evolve throughout the recurrent computation. Our empirical analysis reveals a pronounced late-loop bias in standard low-rank adaptation (LoRA): the shared update provides limited adaptation at early loop positions. This imbalance motivates training shared updates under varying combinations of their applications. Randomly omitting adapter applications alone, however, does not improve task performance; it reduces expected update strength during training while leaving inference unchanged. We introduce Loop Dropout, which couples stochastic masking of adapter applications with inverse-survival rescaling to preserve expected update strength and promote effective adaptation across loops. Extensive experiments demonstrate improved mathematical reasoning across model sizes, adapter ranks and training recipes, with benefits extending to general instruction tuning and code generation. Loop Dropout outperforms existing LoRA variants and adapter regularizers, while further analysis shows stronger early-loop adaptation. Every backbone loop remains active, and inference applies the adapter at all loops using standard LoRA without additional trainable parameters or inference computation. Code is available at https://github.com/NUS-HPC-AI-Lab/loop-dropout .

Figures & tables

Appendix figures & tables25 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Rethinking Adapter Placement: A Dominant Adaptation Module Perspective

    May 7, 2026Suoxin Zhang, Run He, Di Fang +3Low-Rank Adaptation Adapters

  2. Mask the Target: A Plug-and-Play Regularizer Against LoRA Forgetting

    May 28, 2026Runze Xu, Arpit Garg, Hemanth Saratchandran +1Large Language Model AdaptationRapid Adaptation

  3. Balanced LoRA: Removing Parameter Invariance to Accelerate Convergence

    May 29, 2026Valérie Castin, Kimia Nadjahi, Pierre Ablin +1Training-Side Stage-Aware Low-Rank AdaptationLow-Rank Structure