Looped architectures scale computation by reusing the same parameters across recurrent steps, and recent work shows that they substantially improve deep reinforcement learning policies on long-horizon tasks. Since recurrent depth directly controls computation, one may expect looped policies to naturally support elastic inference across recurrent depths. Surprisingly, we find that pretrained looped policies exhibit severe recurrent-depth specialization: reliable decisions are concentrated near the full trained depth, tying deployment computation to this depth even when less computation may suffice. Achieving depth elasticity, i.e., reliable decisions across recurrent depths with adaptive computation at deployment, therefore remains a key challenge. To address this, we propose FlexLoop, a novel post-training framework that converts pretrained fixed-depth looped policies into depth-elastic policies. FlexLoop keeps training on the original RL objective to preserve full-depth capability while performing adjacent-depth policy distillation to progressively transfer decision quality from deeper to shallower recurrent steps. The resulting policy supports reliable inference across recurrent depths and enables state-wise adaptive inference through recurrent-depth consistency. Experiments on 30 online and offline long-horizon goal-conditioned environments show that FlexLoop preserves full-depth performance while making shallower depths effective. Keeping competitive performance, FlexLoop reduces average recurrent depth by up to 43% and achieves up to 1.34× wall-clock speedup in a stress test.
Figures & tables
Figure 1 : Fixed-depth training induces recurrent-depth specialization. Each IRU- K is trained at the K -th recurrent depth. Reliable decisions are concentrated near the full trained depth.
Figure 2 : Framework of FlexLoop. (a) Starting from a depth-specialized looped policy, FlexLoop preserves full-depth capability with the original RL objective while propagating decision quality to shallower depths via adjacent-depth policy distillation. (b) The resulting policy supports reliable decisions across recurrent depths. (c) Adjacent-depth policy discrepancy serves as a stopping signal for state-wise adaptive inference, allocating more computation to harder states.
Figure 3 : (a) FlexLoop reduces adjacent-depth policy discrepancy during post-training and across depths. (b) Smaller δk corresponds to smaller DKL(πk∥πK) on both training and unseen tasks. Marker size denotes the sample fraction in each pooled quantile bin.
Figure 4 : Depth-wise decision performance of pretrained IRU- 5 and FlexLoop on training and unseen tasks. Solid lines denote mean success rates and shading indicates one standard deviation.
Environment
FlexLoop
IRU-5
Avg. Depth
Succ.
Succ.
OGBench
4.19 ± 0.08
0.410 ± 0.020
0.385 ± 0.029
BoxPick (Train)
4.50 ± 0.02
0.989 ± 0.002
0.989 ± 0.003
BoxPick (Unseen)
4.46 ± 0.02
0.778 ± 0.055
0.760 ± 0.043
LightsOut-4x5 (Train)
2.84 ± 0.08
0.747 ± 0.081
0.732 ± 0.076
LightsOut-4x5 (Unseen)
2.93 ± 0.21
0.053 ± 0.031
0.013 ± 0.004
Table 1 : State-wise adaptive inference results across environments. ± indicates one standard deviation.
Model
Success
Avg. Depth
Wall-clock (ms)
Train
IRU-5
0.995±0.001
5
170.1±17.2
FlexLoop
0.993±0.001
3.85±0.04
133.0±7.7
IRU-10
0.997±0.000
10
333.2±36.5
FlexLoop
0.996±0.001
9.35±0.04
319.9±22.2
Unseen
IRU-5
0.511±0.085
5
170.1±17.2
FlexLoop
0.543±0.093
3.80±0.05
127.4±6.3
Table 2 : Wall-clock comparison on BoxPick-Exact-4. ± indicates one standard deviation. Bold numbers indicate the shortest wall-clock latency.
Figure 5 : Ablation studies of FlexLoop. Solid lines denote mean performance, while shaded regions indicate one standard deviation.
Figure 6 : State-wise adaptive recurrent-depth allocation across OGBench tasks. FlexLoop allocates more recurrent computation to states requiring more complex decisions.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Environment
λdepth
BoxPick-Exact-3
1.0
BoxPick-Exact-4
1.0
BoxPick-Generalize-4-1
1.0
BoxPick-Generalize-4-2
1.0
LightsOut
0.05
Scene
0.05
Appendix
Table 3 : Adjacent-depth policy distillation weight used for each environment.
Figure 7 : Depth-wise decision performance of pretrained IRU- 5 and FlexLoop on unseen tasks. Solid lines denote mean success rates and shading indicates one standard deviation. FlexLoop improves shallow-depth decision elasticity while maintaining out-of-distribution generalization.
Figure 8 : Depth-wise decision performance during IRU pretraining and FlexLoop post-training. Solid lines denote mean success rates and shading indicates one standard deviation. FlexLoop rapidly improves shallow-depth performance while preserving the full-depth capability of pretrained policies.
Figure 9 : Depth-wise decision performance on training tasks before and after FlexLoop post-training. Markers denote mean success rates and shading indicates one standard deviation.
Figure 10 : Depth-wise decision performance on unseen tasks before and after FlexLoop post-training. Markers denote mean success rates and shading indicates one standard deviation.
Task
FlexLoop
IRU-5
Avg. Depth
Succ.
Succ.
Scene
4.66 ± 0.06
0.496 ± 0.064
0.489 ± 0.087
Cube
4.75 ± 0.03
0.569 ± 0.056
0.492 ± 0.114
Puzzle
3.45 ± 0.03
0.296 ± 0.029
0.293 ± 0.01
AntMaze
4.80 ± 0.03
0.292 ± 0.059
0.253 ± 0.034
PointMaze
3.28 ± 0.30
0.395 ± 0.004
0.396 ± 0.003
Appendix
Table 4 : Full state-wise adaptive inference results on all environments. ± indicates one standard deviation.
Figure 11 : Pareto trade-offs between task performance and average recurrent depth under different adaptive inference thresholds. Markers denote mean success rates and error bars indicate one standard deviation.
Figure 12 : Selected trajectory segments from a single episode in the Scene environment. FlexLoop adaptively uses fewer recurrent steps for simple, monotonic motions and increases computation at challenging decision points, such as goal switching.
Figure 13 : Depth-wise performance evolution of pretrained IRU-5 and FlexLoop from scratch on both training and unseen tasks. Solid lines denote mean success rates and shading indicates one standard deviation.
Figure 14 : Elasticity and Pareto curves of pretrained IRU-5 and FlexLoop from scratch on both training and unseen tasks. Markers denote mean success rates and the shading and error bars indicate one standard deviation.
Figure 15 : Elasticity and Pareto Curves of the policies with different pretraining recurrent depths before and after FlexLoop post-training. Markers denote mean success rates and shading indicates one standard deviation.
Humans and machines often solve harder problems by spending more time on computation. In deep learning, looped models implement this idea during inference by recurrently updating a hidden state. In practice, however, their training backpropagates through only one or a few updates, making it hard to train early updates to support future ones. We propose looped flows, an approach that sidesteps this issue by training the recurrence with local denoising objectives. By imposing temporal association across denoising objectives through progressively decreasing noise levels and shared noise, the model is incentivized to learn recurrent states that transfer useful computation over time, even when gradients cover only a few updates. We then formulate inference as integrating the velocity of a probability flow parameterized by the learned denoiser, coupled with recurrent states. This allows solving harder problems by spending more computation through a finer temporal grid and enables multiple valid predictions from different initial noise samples. Across six reasoning benchmarks including two multi-solution benchmarks, looped flows outperform prior state-of-the-art looped models overall, achieving 58.8% test accuracy on ARC-AGI-1 and 12.2% on ARC-AGI-2.
How does the amount of compute available to a reinforcement learning (RL) policy affect its learning? Can policies using a fixed amount of parameters, still benefit from additional compute? The standard RL framework does not provide a language to answer these questions formally. Empirically, deep RL policies are often parameterized as neural networks with static architectures, conflating the amount of compute and the number of parameters. In this paper, we formalize compute bounded policies and prove that policies which use more compute can solve problems and generalize to longer-horizon tasks that are outside the scope of policies with less compute. Building on prior work in algorithmic learning and model-free planning, we propose a minimal architecture that can use a variable amount of compute. Our experiments complement our theory. On a set 31 different tasks spanning online and offline RL, we show that (1) this architecture achieves stronger performance simply by using more compute, and (2) stronger generalization on longer-horizon test tasks compared to standard feedforward networks or deep residual network using up to 5 times more parameters.
Raj Ghugare, Michał Bortkiewicz, Alicja Ziarko +1
Department of Computer Science, Princeton University · Warsaw University of Technology · University of Warsaw +2
Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters. This reuse changes the residual-scaling problem: in an untied Transformer, each residual branch receives and applies its own parameter update, whereas in a looped Transformer one shared update aggregates gradients from repeated visits and is read back by those same visits in the next linearized forward pass. We formalize this tied-depth effect through a first-order perturbation bound controlled by a visit-alignment coefficient κR. The bound recovers the DeepNorm exponent when visits decorrelate, but in the conservative aligned regime it requires the exponent to increase from 1/4 to 1/2 as loop count grows at fixed physical depth. The resulting method, \textbf{DeepLoop}, keeps the Post-LN DeepNorm architecture and sets α=(2N)1/2 and β=(8N)−1/2 for unrolled depth N. On GPT-style looped language models at GPT-2 small and GPT-2 medium scale, DeepLoop is neutral when no physical block is revisited and improves validation loss and downstream accuracy once recurrent depth is activated. These results show that stable recurrent depth requires residual scaling rules that account for parameter visits, not only nominal layer count.
Shuzhen Li, Yifan Zhang, Jiacheng Guo +2
Princeton University · University of California, Los Angeles