cs.LGSep 28, 2026

FlexLoop: Depth-Elastic Looped Policies for Adaptive Test-Time Computation in Deep RL

Authors: Xun Wang, Ruishuo Chen, Yu Chen, Zhuoran Li, Longbo Huang

Organizations: Institute for Interdisciplinary Information Sciences, Tsinghua University

Abstract

Looped architectures scale computation by reusing the same parameters across recurrent steps, and recent work shows that they substantially improve deep reinforcement learning policies on long-horizon tasks. Since recurrent depth directly controls computation, one may expect looped policies to naturally support elastic inference across recurrent depths. Surprisingly, we find that pretrained looped policies exhibit severe recurrent-depth specialization: reliable decisions are concentrated near the full trained depth, tying deployment computation to this depth even when less computation may suffice. Achieving depth elasticity, i.e., reliable decisions across recurrent depths with adaptive computation at deployment, therefore remains a key challenge. To address this, we propose FlexLoop, a novel post-training framework that converts pretrained fixed-depth looped policies into depth-elastic policies. FlexLoop keeps training on the original RL objective to preserve full-depth capability while performing adjacent-depth policy distillation to progressively transfer decision quality from deeper to shallower recurrent steps. The resulting policy supports reliable inference across recurrent depths and enables state-wise adaptive inference through recurrent-depth consistency. Experiments on 3030 online and offline long-horizon goal-conditioned environments show that FlexLoop preserves full-depth performance while making shallower depths effective. Keeping competitive performance, FlexLoop reduces average recurrent depth by up to 43%\bf{43\%} and achieves up to 1.34×\bf{1.34\times} wall-clock speedup in a stress test.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Thinking with Looped Flows

    Sep 12, 2026Ayhan Suleymanzade, Chanhyuk Lee, Floor Eijkelboom +3Deep LearningRecurrent State

  2. On the Role of Computation in Reinforcement Learning

    Feb 5, 2026Raj Ghugare, Michał Bortkiewicz, Alicja Ziarko +1Long-Horizon Task PlanningRole

  3. DeepLoop: Depth Scaling for Looped Transformers

    Jul 15, 2026Shuzhen Li, Yifan Zhang, Jiacheng Guo +2Transformer Residual StreamsTransformer Architectures