Recursive reasoning models apply a small shared Transformer block many times to refine a latent state. This gives them large effective depth with few parameters and makes them strong on algorithmic tasks. Such compact solvers are natural candidates for tools that an LLM can call on narrow algorithmic subproblems. However, existing models such as HRM, TRM and URM differ in architecture, gradient propagation and training procedure simultaneously. This makes it hard to tell what drives their performance, and their optimization is still poorly understood and often unstable. In this work we address both of these gaps. First, we study these questions under a unified experimental pipeline spanning six algorithmic domains. Individual controlled ablations are performed on representative domains, while the resulting recipe is evaluated across the full suite. The study reveals a surprisingly simple recipe for stable and generalizable recursive reasoning: an intermediate gradient horizon, large physical batches and controlled updates of the recurrent state. An explicit hierarchical architecture is not needed. Second, we combine these findings into a stable 13.6M-parameter model that achieves the strongest overall performance among the evaluated recursive baselines, with particularly large gains on out-of-distribution generalization. It raises Arithmetic OOD accuracy to 71.2%, from 36.2% for the strongest baseline, while reaching 98.41% on Sudoku and 59.5% pass@2 on ARC-AGI-1. Our results show that, within the recursive architectures studied here, performance depends strongly on how recurrence is optimized and stabilized. More broadly, it shows how AI systems can be improved by optimizing their components one at a time.
Figures & tables
Figure 1: Recursive reasoning performance is determined by both recurrent computation and its optimization. We systematically study the factors that govern recursive models: gradient propagation through the recurrent trajectory, recurrent-state stabilization, architectural choices, and test-time computation. Our final 13.6M-parameter model combines the most effective components and achieves strong performance across algorithmic domains while improving out-of-distribution generalization.
Figure 2: Evaluation domains used in our study. Sudoku, Maze, and ARC-AGI follow established recursive-reasoning benchmarks, while Game of Life and Arithmetic provide controlled out-of-distribution evaluation.
Figure 3: Overview of the final recursive reasoning architecture. A shared Transformer block repeatedly refines latent recurrent states. Controlled state updates and truncated gradient propagation stabilize the recursive trajectory.
Setup
HRM
TRM
URM
Ours
Parameters
27M
7M
7M
13.6M
Modules
2 ( H+L )
1 shared
1 shared
1 shared
Layers
4+4
2
2*
4
Hidden size
512
512
512
512
Cycles (H,L)
(2,2)
(3,6)
(3,6)
(4,2)
Gradient path
Last H + last L
Last H + all L
Last H + all L
(KL,KH)=(2,2)
Table 1: Comparison of recursive reasoning recipes.
Metric
Dense
Qwen3-30B-A3B
Qwen3.6-27B
HRM
TRM
URM
Ours
Arithmetic ID
78.50
0.10
0.00
92.00
98.80
95.60
96.34
Arithmetic OOD
7.49
0.00
0.00
13.80
36.20
28.50
71.16
Game of Life ID
8.17
0.00
0.00
5.00
7.11
55.42
66.08
Game of Life OOD
7.54
0.00
0.00
0.30
7.14
55.26
65.80
Maze
8.00
0.20
0.00
71.10
80.00
80.20
84.70
Sudoku
18.30
0.00
0.00
50.00
87.40
77.60
98.41
Table 2: Exact accuracy on the evaluation domains, in percent, from EMA weights. Arithmetic is split into the in-distribution and out-of-distribution target buckets of Section 3 ; Game of Life is reported on its in-range set and as the mean over its out-of-distribution sets. ARC-AGI is evaluated with pass@1 and pass@2. HRM, TRM, URM, and the dense control are retrained in our codebase; general-purpose models are evaluated from released checkpoints without fine-tuning. Best values are in bold.
Figure 4: Exact accuracy on Arithmetic as a function of the differentiated gradient horizon. Forward recursion is fixed at Lcycles=2 and Hcycles=4 ; only (KL,KH) varies. Left: in-distribution accuracy. Right: out-of-distribution accuracy.
Figure 5: Effect of physical batch size versus gradient accumulation at a fixed effective batch size of 8192 on Game of Life. Configurations correspond to physical batch sizes of 8192, 4096, and 2048 with 1, 2, and 4 accumulation steps, respectively. Left: ID and OOD accuracy. Right: mean pre-clipping gradient norm.
Configuration
GoL ID
GoL OOD
Sudoku
Base
61.95±3.32
63.87±2.90
97.29±1.06
+ gate
64.24±2.59
65.89±2.77
98.41±0.25
+ grad clip
62.41±1.56
63.87±1.54
98.50±0.20
+ dropout
61.82±0.34
62.09±0.41
98.36±0.25
+ relative noise (final)
66.08±0.11
65.80±0.13
98.41±0.21
Table 3: Cumulative stabilization ablation on Game of Life and Sudoku. Post-cycle normalization and EMA are fixed in all configurations. Values report mean ± standard deviation over three seeds. Best values are in bold and second-best values are underlined.
Figure 6: Architectural and test-time scaling ablations. (a) Effect of sharing parameters between recurrent timescales. (b) Effect of additional recurrent computation at inference time.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Domain
Input structure
Primary capability
Controlled OOD
Sudoku
9×9 grid
Constraint satisfaction
–
Maze
30×30 grid
Search and planning
–
Game of Life
Variable-size binary grid
Iterative dynamics
Patterns, horizon
Arithmetic
RPN expression
Symbolic composition / search
Operands, value range
ARC-AGI-1
Grid transformations
Abstract rule induction
–
ARC-AGI-2
Grid transformations
Abstract rule induction
–
Appendix
Table 4: Summary of the six evaluation domains. Game of Life and Arithmetic provide explicit controlled distribution shifts; the remaining domains are retained primarily for comparability with prior recursive reasoning models.
Domain
dcore
dH=dL
Conv.
η
zL -gate
Arithmetic
.025
.010
no
.005
yes
Sudoku
.010
.000
no
.005
yes
Game of Life
.100
.000
no
.005
yes
Maze
.010
.000
no
.005
yes
ARC-AGI-1
.025
.025
yes
.003
yes
ARC-AGI-2
.025
.025
yes
.005
yes
Appendix
Table 5: Domain-specific recurrent regularization. dcore denotes dropout in the embedding, attention, residual, and feed-forward paths, while dH and dL denote dropout applied to the recurrent states zH and zL . η is the relative state and embedding noise scale. “Conv.” indicates convolutional input mixing. Controlled zL updates use τ=0.7 .
Domain
Physical batch
Epochs
Eval. interval
Peak LR
Min. LR ratio
β2
Weight decay
Puzzle LR
Arithmetic
4096
2000
50
5×10−4
.01
.95
1.0
.005
Sudoku
4096
2000
50
1×10−4
1.0
.95
0.1
.010
Game of Life
8192
2000
50
1×10−4
.10
.95
0.1
.010
Maze
1024
54000
500
1×10−4
.01
.995
0.1
.010
ARC-AGI-1
768
300000
10000
1×10−4
1.0
.95
0.1
.010
ARC-AGI-2
768
300000
10000
1×10−4
1.0
.95
0.1
.010
Appendix
Table 6: Domain-specific optimization configurations. All runs use 2,000 learning-rate warmup steps, β1=0.9 , global gradient clipping at norm 1, and EMA weights. Puzzle embeddings use a separate learning rate (“Puzzle LR”) and the same weight decay as the remaining parameters.
Model
Layers (H,L)
Cycles (H,L)
Relative compute
Dense Transformer
8,−
–
1
HRM
4,4
(2,2)
3T
TRM
2,2
(3,6)
5.25T
URM
2,2
(3,6)
5.25T
Ours
4,4
(4,2)
6T
Appendix
Table 7: Inference cost of one answer relative to a single forward pass of the non-recurrent dense control. T denotes the number of ACT steps.
Figure 7: Effect of the gradient-path length on Arithmetic. Top: final EMA best_full exact accuracy as a function of the number of high-level cycles retained in the backward graph, for one or two retained low-level cycles. Bottom: OOD accuracy over training. The selected (2,2) model obtains the highest final OOD score.
gL
gH
Final ID
Final OOD
Best OOD
Best update
1
1
97.35
13.87
19.67
87,920
1
2
99.54
7.93
20.91
87,920
1
3
99.60
8.80
17.88
428,610
1
4
99.31
65.53
68.54
428,610
2
1
98.12
23.71
28.61
351,680
2
2
96.34
71.16
71.16
428,610
Appendix
Table 8: EMA best_full exact accuracy under gradient truncation. Accuracies are percentages. Final values are measured at 439,600 updates. “Best OOD” is the maximum over the 40 recorded checkpoints and uses the OOD split for checkpoint selection.
Figure 8: Single-seed Game of Life training trajectory. Left: EMA best_full exact accuracy on ID and five OOD splits. Right: the unweighted OOD mean for online and EMA parameters; shading spans the minimum and maximum EMA OOD split. Accuracy peaks at epoch 1150 and declines before the final checkpoint.
Split
Online final
Online peak
EMA final
EMA peak
EMA token acc.
Best forced step
ID
63.70
65.99
64.01
65.98
75.81
63.75 (4)
OOD +1 step
63.71
65.73
64.04
65.94
75.83
63.79 (4)
OOD +2 steps
63.65
65.73
63.94
65.83
75.92
63.67 (3)
OOD +3 steps
63.02
65.22
63.35
65.37
75.45
63.03 (3)
OOD +10 steps
63.46
65.94
64.04
65.87
75.92
63.76 (4)
OOD unseen patterns
63.12
65.34
63.46
65.36
75.69
63.12 (5)
Appendix
Table 9: Final and peak validation performance. Values are percentages. Exact columns use the logged best_full exact-accuracy series; token accuracy is the EMA accuracy series at the final checkpoint. All EMA peaks occur at 211,140 updates (epoch 1150). “Best forced step” reports the highest single-step exact score at the final checkpoint, followed by its logged ACT step in parentheses.
Figure 9: Final and peak EMA best_full quality and recurrent-depth profile. Left: final and peak best_full exact accuracy for every evaluation split. Right: EMA single-step exact accuracy at the final complete checkpoint. Most of the single-step gain occurs between forced ACT steps 1 and 2; all six curves then remain in a narrow band through step 24.
Figure 10: Exact validation accuracy during joint four-domain training. Online and EMA parameters are evaluated at six checkpoints. Arithmetic improves throughout training, but Game of Life remains below 0.1% exact accuracy and neither Maze nor Sudoku produces a single exact solution. The panels use different vertical scales.
Domain
Final token acc.
Best exact acc.
Final exact acc.
Arithmetic
63.97
28.22
28.22
Game of Life
51.68
0.09
0.06
Maze
87.50
0.00
0.00
Sudoku
20.21
0.00
0.00
Appendix
Table 10: EMA validation metrics for joint training. Values are percentages. “Best” is selected over the six checkpoints; final metrics are measured at 194,562 updates. Token accuracy can be high even when the full structured output is always incorrect.
Recursive models create computational depth through parameter reuse, offering a parameter-efficient alternative to increasing model size. However, most recursive models repeatedly apply one learned transformation or a prescribed sequence of transformations, restricting computation to a fixed layer order. We introduce the Random Recursive Model (RRM), which maintains a pool of L learned layers and performs T recursive steps by sampling one layer independently with replacement for each example and step. This enables flexible layer reuse while retaining the parameter efficiency of recurrence. We evaluate RRM on challenging reasoning tasks, where it matches or exceeds the baselines, often with 50-75 % fewer parameters. RRM can vary its depth at inference, including beyond that seen during training, without retraining or adding parameters, improving tasks that benefit from deeper iterative computation. RRM also supports Monte Carlo inference and probabilistic test-time scaling, both of which improve performance without retraining. These insights may open new directions in neural network architecture design.
Jama Hussein Mohamud, Mirco Ravanelli
Mila – Quebec AI Institute · Université de Montréal · Concordia University
Hard symbolic-reasoning tasks such as Sudoku, maze pathfinding, and ARC remain challenging for LLMs due to their fixed-depth autoregressive reasoning, which limits systematic search, refinement, and backtracking. While recursive models such as Hierarchical Reasoning Model (HRM) and Tiny Recursive Model (TRM) address this limitation through iterative latent-state refinement, they are typically task-specific and do not leverage pretrained language priors. We propose R-Qwen, a recursive reasoning framework built upon a pretrained Qwen backbone. R-Qwen repeatedly refines a candidate solution through programmatic self-recursion and deep supervision, combining the structured iterative computation of recursive models with the linguistic and reasoning priors of pretrained LLMs. We further adapt Hierarchical Supervision Weighting (HSW) to autoregressive models by exponentially weighting losses across recursive steps. HSW reduces gradient variance by at least 50%, improves the signal-to-noise ratio of stochastic gradients, and accelerates convergence. Across eight challenging benchmarks, R-Qwen consistently outperforms prior recursive reasoning models and substantially larger LLMs while using a comparable number of trainable parameters. Notably, on ARC-AGI dataset, our model achieves a 27.6% improvement over the baseline, highlighting the effectiveness of recursive refinement for general symbolic reasoning. These results suggest that recursive reasoning mechanisms and pretrained language model priors are complementary approaches for improving symbolic puzzle-solving. Code and models will be released after acceptance.
Modern language models reason within bounded context, an inherent constraint that poses a fundamental barrier to long-horizon reasoning. We identify recursion as a core principle for overcoming this barrier, and propose recursive models as a minimal realization, where the model can recursively invoke itself to solve subtasks in isolated contexts. We prove that any computable problem admits a recursive decomposition of reasoning in which each subtask requires only exponentially smaller active context than standard autoregressive models; this strictly surpasses any context management approach confined to a single sequence, such as summarization. We further generalize our framework to modern agentic systems with arbitrary context processing and control flows, and prove that recursive models can achieve optimal power within this broader class. Experimentally, we test two settings: fine-tuning a pretrained base model for recursive SAT solving, and training a small model from scratch on Go traces generated by exact game-tree search. Both show improved long-horizon accuracy with small active contexts.