Scaling to Tens of Thousands of Test-Time Iterations with Loop-Native Attention Residuals
Authors: Pengxiang Li, Dilxat Muhtar, Di He, Guinan Su, Lu Yin, Shiwei Liu
Organizations: The Hong Kong Polytechnic University · Independent Researcher · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences · Peng Cheng Laboratory · University of Chinese Academy of Sciences · ELLIS Institute Tübingen · Max Planck Institute for Intelligent Systems · Tübingen AI Center · Shenzhen University of Advanced Technology
In this paper, we argue that looped Transformers need their own residual connections to prevent performance degradation as the number of iterations grows. We observe that increasing loop iterations can reduce reasoning accuracy: noisy state updates overwrite correct intermediate deductions and even undo completed solutions. This leaves subsequent iterations to recover lost information from an already degraded representation: once an error arises in an earlier loop, often as a result of long-range propagation through the recurrence, later loops find it difficult to correct. In this paper, we introduce InfiLoop, a loop-native residual connection that learns which past computations to retain and how much to accept from each new update. InfiLoop combines content-based weighting with learned temporal decay to maintain a running summary of recurrent states. An exact streaming recurrence keeps its persistent aggregation memory constant as the loop count grows. The resulting adaptive update suppresses unreliable proposals and preserves useful intermediate states. Across extensive reasoning tasks, a 7M-parameter InfiLoop model outperforms existing recursive architectures, reaching 97.9% exact accuracy on Sudoku-Extreme, and 13.6% pass@2 on ARC-AGI-2. Notably, on Sudoku-Extreme, InfiLoop continues to improve with test-time looping beyond 20,000 effective steps, showing that added depth translates directly into stronger reasoning. Our code is available at https://github.com/pixeli99/InfiLoop.
Figures & tables
Figure 1: Carry-last keeps changing its state after solving and loses solutions. (a) Exact accuracy against effective depth, the number of executed Transformer layers, on Sudoku-Extreme. Each outer update executes 26 layers. TRM peaks at 468 layers and then loses accuracy as depth grows. (b) Relative state change after the update at which each model first solves a puzzle, for the puzzles that both models solve within 36 updates. TRM settles briefly and then moves its state again. (c) Share of those puzzles whose solution is still correct.
Figure 2: Connecting successive applications of a shared block. (a) TRM uses the carry-last rule: the candidate zt becomes the next recurrent state, ht=zt . (b) InfiLoop combines earlier candidates into the state ht passed to the next block application, using content weights and exponential temporal decay. Colors identify candidates from different iterations; the illustrated contributions are schematic. The same block parameters are reused at every recurrent step.
Puzzles (exact acc.)
ARC-AGI (pass@2)
Method
Params
Sudoku-Ext.
Maze-Hard
ARC-1
ARC-2
Large language models
DeepSeek R1
671B
0.0
0.0
15.8
1.3
Claude 3.7 Sonnet
–
0.0
0.0
28.6
0.7
o3-mini-high
–
0.0
0.0
34.5
3.0
Gemini 2.5 Pro (32K)
–
–
–
37.0
4.9
Table 1: Reasoning benchmark results (%) for complete systems under their respective inference procedures. Sudoku and Maze use exact solution accuracy. ARC uses pass@2. Best results are bold, and the best among the listed 7M-parameter models are underlined. Results for DeepSeek R1 ( DeepSeek-AI, 2025 ) , Claude 3.7 Sonnet ( Anthropic, 2025 ) , o3-mini ( OpenAI, 2025 ) , and Gemini 2.5 Pro ( Comanici et al., 2025 ) follow Jolicoeur-Martineau (2025) ; URM and EqR parameter counts follow Movahedi et al. (2026) . ‡ URM reports ARC pass@1. † Reproduced by us: TRM retrained on Sudoku-Extreme with the authors’ published recipe, and the released ARC-AGI-2 verification checkpoint evaluated under our protocol.
Figure 3: Test-time scaling and accuracy across puzzle difficulty. Left: Exact accuracy against executed Transformer layers (logarithmic axis), with all three models evaluated to 24,960 layers. InfiLoop reaches 90.9% at 1,872 layers, exceeding FPRM’s 90.1% at 3,744 layers with half the layer evaluations. Right: Exact accuracy by empty-cell count at matched 1,872 layers.
Figure 4: Learned adaptive updates reflect solution milestones and board correctness. (a) Mean adaptive step size γt aligned to the update at which the puzzle is first solved (for puzzles solved at update 13 or later). The step size remains suppressed ( γt≈0.40 ) during search, peaks at 0.91 when the puzzle is completed, and settles to the equal-weight baseline 1−β=0.56 thereafter. (b) Median clipped content score lnet=clip(st,−8,8) for block outputs produced after update 12, plotted against the number of incorrect cells in the decoded board. The shaded band denotes the interquartile range. High weights are assigned exclusively to valid configurations without explicit correctness supervision.
Figure 5: State stability and solution retention under recurrent state perturbation. At outer update 36, a relative perturbation ϵ∈[10−4,0.3] is injected into the recurrent state, followed by 36 additional updates. (a) Proportion of puzzles solved at update 36 that remain solved at update 72 as a function of perturbation magnitude ϵ . The horizontal dashed line marks the unperturbed baseline. (b) Evolution of the median relative Euclidean distance between the perturbed and unperturbed states, log10(∥h36+k′−h36+k∥/∥h36+k∥) , across subsequent updates k . Different line shades denote perturbation magnitudes ϵ . While perturbations in TRM quickly diverge to O(1) relative distance, InfiLoop constrains trajectory drift and retains all solved states.
Connection placement across budgets
Configuration
H3L6
H3L12
H3L96
H32L6
Inner axis alone
72.6
75.2
77.2
79.4
Outer axis alone
75.1
76.3
76.2
80.0
Both axes (ours)
82.1
85.3
87.4
89.7
Table 2: Component and initialization ablations on Sudoku-Extreme. Exact solution accuracy (%) on 7,680 test puzzles. Top: placement of the connection across recurrent axes at training depth H3L6 and three larger budgets. Bottom: temporal decay on the two-axis model. The left group fixes β during training. The right group learns β from the initialization β0 .
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Qualitative comparison on ARC-AGI-2. Each row shows a demonstration input–output pair, a test input, and the predictions of InfiLoop and TRM. The top and bottom rows correspond to tasks 5961cc34 and 80a900e0 , respectively. As in the pass@2 evaluation, we show the best of each model’s two submitted attempts. InfiLoop matches the target exactly in both cases; TRM makes 30 and 14 cell errors, respectively.
InfiLoop
q=0 in training
q=0 at test time
β=0 at test time
90.9
84.1
84.8
1.5
Appendix
Table 3: Exact accuracy (%) on Sudoku-Extreme at H6L12 , S=12 , or 1,872 layers.
Looped architectures provide an inductive bias toward learning step-by-step procedures for tasks that require compositional reasoning. The number of effective layers reached by looping determines the quality of the solution these models find. Like deep architectures, looped architectures are prone to a signal propagation problem induced by depth as the halting decision is postponed. In this paper, we address this signal propagation issue using pre-norm layers and residual scaling. Building on these architectural modifications, we propose FPRM, a Transformer-based Fixed-Point Reasoning Model that uses fixed-point convergence as an end-to-end halting mechanism in a looped architecture. We show that fixed-point halting allows FPRM to adapt its compute to task difficulty. FPRM is effective on common reasoning benchmarks, namely Sudoku, Maze, state-tracking, and ARC-AGI.
Sajad Movahedi, Vera Milovanović, Shlomo Libo Feigin +5
ELLIS Institute Tübingen, Max Planck Institute for Intelligent Systems, Tübingen AI Center · ETH Zurich · Swiss Institute of Bioinformatics +2
Humans and machines often solve harder problems by spending more time on computation. In deep learning, looped models implement this idea during inference by recurrently updating a hidden state. In practice, however, their training backpropagates through only one or a few updates, making it hard to train early updates to support future ones. We propose looped flows, an approach that sidesteps this issue by training the recurrence with local denoising objectives. By imposing temporal association across denoising objectives through progressively decreasing noise levels and shared noise, the model is incentivized to learn recurrent states that transfer useful computation over time, even when gradients cover only a few updates. We then formulate inference as integrating the velocity of a probability flow parameterized by the learned denoiser, coupled with recurrent states. This allows solving harder problems by spending more computation through a finer temporal grid and enables multiple valid predictions from different initial noise samples. Across six reasoning benchmarks including two multi-solution benchmarks, looped flows outperform prior state-of-the-art looped models overall, achieving 58.8% test accuracy on ARC-AGI-1 and 12.2% on ARC-AGI-2.
Looped language models increase effective depth by repeatedly applying a shared block of layers, but existing large-scale recipes require multi-stage training over trillions of tokens, while the benefits of recurrence remain difficult to separate from differences in data and training. In this work, we establish practical training recipes for looped language models, with three main results. (1) We develop a compute-efficient from-scratch pipeline that reduces the training budget from 7.7T tokens in Ouro to 310B tokens while retaining strong reasoning performance. Pretraining followed by high-quality mid-training, together with learning-rate warmup and stronger exit-gate regularization, enables stable recurrent training without prior multi-stage schedules. (2) Under controlled comparisons, our 1.4B LoopLM outperforms a parameter-matched dense model trained on the same data and token budget on all 12 evaluated benchmarks, including +14 points on GSM8K, +10 on MATH, and +22 on DROP. At matched inference compute, it approaches a 3.9B dense model on mathematical reasoning and reading comprehension while using only 36% as many parameters. (3) We introduce a minimal recipe for converting pretrained dense models into looped ones: a single learned input-mixing scalar and a smoothed exit loss, with no step-specific parameters. Applied to Qwen3-1.7B-Base, Looped Qwen improves over an identically continued dense baseline on every evaluated benchmark across two data regimes, with statistically clear gains on GSM8K, MATH, and MMLU-Pro on the curated mixture. Together, these results make looped language models substantially cheaper to train from scratch and practical to introduce into existing pretrained checkpoints, while isolating the gains due to recurrence itself.
Andrei Marchenko, Viacheslav Bezrukov, Oleg Kashurin +5