cs.AIOct 6, 2026

Enabling Dynamic Computation in Looped LMs

Authors: Aayush Mishra, Arnau Padrés Masdemont, Victor Conchello Vendrell, Jordi Ros Giralt, Arash Behboodi, Fabio Valerio Massoli

Organizations: Qualcomm AI Research · Johns Hopkins University

Abstract

Looped LMs are parameter efficient and promise dynamic computation (saving memory and FLOPs on easy tokens). However, state-of-the-art open Looped LMs trained with this dynamic computation capability (Ouro models) do not realize it in practice as each loop iteration (depth) requires its own level of KV-cache, necessitating all loop computations. Moreover, Ouro's early-exit prior is enforced on each token equally, which results in static lower-depth like processing of all tokens regardless of difficulty. In this work, we propose a simple "best-available" KV caching strategy that works out-of-the-box, creating a new frontier in the performance vs depth space. Our approach enables up to 30% reduction in FLOPs and KV memory while retaining full-depth performance, showing the true flexibility of Looped LMs. Furthermore, training looped LMs with awareness about this KV caching strategy improves performance and efficiency. Finally, we apply a small but effective fix to the early-exit prior enforcement objective that makes tokens exit at truly heterogeneous depths based on effort. Our findings are validated on Ouro models as well as smaller looped LMs pre-trained from scratch.

Figures & tables

Appendix figures & tables31 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Closing the Loop: Practical Training Recipes for Looped Language Models

    Sep 30, 2026Andrei Marchenko, Viacheslav Bezrukov, Oleg Kashurin +5Recurrent ModelLoop

  2. Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching

    Aug 10, 2026Kristian Schwethelm, Daniel Rueckert, Georgios KaissisEfficient Inference

  3. FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates

    Sep 24, 2026Wanqi Yang, Shiwei LiuBlock Sparse Flash AttentionTransformer Architectures