Recursive models create computational depth through parameter reuse, offering a parameter-efficient alternative to increasing model size. However, most recursive models repeatedly apply one learned transformation or a prescribed sequence of transformations, restricting computation to a fixed layer order. We introduce the Random Recursive Model (RRM), which maintains a pool of L learned layers and performs T recursive steps by sampling one layer independently with replacement for each example and step. This enables flexible layer reuse while retaining the parameter efficiency of recurrence. We evaluate RRM on challenging reasoning tasks, where it matches or exceeds the baselines, often with 50-75 % fewer parameters. RRM can vary its depth at inference, including beyond that seen during training, without retraining or adding parameters, improving tasks that benefit from deeper iterative computation. RRM also supports Monte Carlo inference and probabilistic test-time scaling, both of which improve performance without retraining. These insights may open new directions in neural network architecture design.
Figures & tables
Figure 1: Fixed-order depth, fixed-recursive depth, and RRM. Four applications under the three depth constructions. Fixed-order depth applies each layer once in its prescribed position. Fixed-recursive depth repeatedly applies a prescribed sequence of learned functions. Here, the repeated sequence contains one function ( L=1 ); more generally, it may contain multiple layers and be surrounded by fixed-order layers. RRM samples from the complete set of learned layers at every step, allowing both reordering and reuse. Colors identify shared parameters.
Figure 2: CIFAR-10 representations across recursive depth. We take the class-token representations from the RRM trained with L=2 and Ttrain=16 and plot them at Ttest∈{8,12,16,20,24} , using one sampled layer sequence per image. Each panel is a t-SNE projection of the same test images, colored by class. Both the k -NN classification accuracy and the t-SNE projections show that the representations become more class-specific as the number of recursive applications increases.
Model
L
Ttrain
Params.
M
Accuracy (%)
Dense
8
8
3.20M
1
92.520±0.000
RRM
2
8
0.824M
1
91.403±0.078
RRM
2
16
0.824M
1
91.670±0.155
RRM
2
16
0.824M
16
92.337±0.049
RRM
2
8
0.824M
16
92.597±0.110
Table 1: CIFAR-10 top-1 accuracy. M denotes Monte Carlo inference samples.
Table 4
Model
L
T
Params.
Accuracy (%)
HRM
8
24
27M
55.00
TRM
2
42
5M
87.63
RRM
2
21
5M
87.22±0.03
TRM+RRM
2
42
5M
89.17±0.02
RRM †
2
32
5M
89.21±0.03
Table 4: Sudoku-Extreme exact-grid accuracy. HRM, TRM, and TRM+RRM use Nsup=16 , while RRM uses Nsup=32 . † evaluates the same RRM with Ttest=32 without retraining. The TRM row is our reproduction; the published result is 87.40% ( Jolicoeur-Martineau, 2025 ) .
Model
L
T
T×Nsup
Params.
Pass@2 (%)
HRM
8
24
384
27M
40.30
TRM-Att
2
30
480
7M
44.60
RRM †
2
15
30
7M
45.79±0.29
RRM
2
15
480
7M
47.38±0.25
Table 5: ARC-AGI-1 Pass@2 accuracy. Baseline results are from Jolicoeur-Martineau (2025) . † evaluates RRM with Nsup=2 instead of the 32 used during training.
Figure 3: Top: Parameter efficiency across four tasks. RRM matches or outperforms the corresponding baselines with substantially fewer parameters in (a–c) and at a comparable parameter count in (d). Sudoku LM denotes the autoregressive Sudoku task in Section 5.2 . Bottom: Inference compute efficiency across four reasoning tasks. FLOPs are measured per evaluation example and include all applications across supervision steps when deep supervision is used. Overall, RRM consistently improves parameter efficiency and achieves competitive or better compute efficiency on three of the four reasoning tasks, with the largest gains on Maze-Hard and ARC-AGI-1.
Figure 4: Effect of the number of reusable layers. RRM remains effective with L=1 , while L=2 matches or exceeds the baselines on both tasks.
Figure 5: Top: Test-time scaling of RRM without retraining or adding parameters. Vertical dotted lines mark Ttrain . Horizontal dashed lines show the corresponding baselines. Additional test-time applications leave performance in (a,b) largely unchanged but improve it in (c,d). Bottom: Train-test depth generalization. Rows indicate Ttrain and columns indicate Ttest ; outlined cells mark Ttrain=Ttest . Across both tasks, performance decreases as Ttest moves farther below Ttrain . When Ttest>Ttrain , Maze-Hard improves or plateaus, whereas CIFAR-10 plateaus or degrades depending on Ttrain .
Model
Standard (%)
PTRM (%)
TRM
87.20
98.43±0.12
TRM+RRM
88.57±0.25
98.83±0.06
RRM
87.20±0.26
99.50±0.10
Table 6: Sudoku-Extreme accuracy with standard inference and PTRM test-time scaling.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Model
L
T
M
Accuracy (%)
Dense
8
8
1
92.52
Learned policy
8
7.99
1
85.07
RRM
2
8
1
91.40
RRM
2
8
16
92.60
RRM
8
32
1
93.67
Appendix
Table 7: CIFAR-10 learned-ordering ablation. The learned policy’s reported T is the mean number of applications. M denotes Monte Carlo inference samples.
Figure 6: CIFAR-10 training-depth and test-time-scaling experiments. All RRMs use L=2 . Vertical dotted lines mark Ttrain , and horizontal dashed lines show the dense eight-layer baseline.
Model
Tdetach
Ttrain
Ttest
Nsup
Total
Pass@2 (%)
TRM-Att
–
30
30
16
480
44.60
RRM
0
8
8
32
256
42.67±0.14
RRM
0
8
15
32
480
42.79±0.26
RRM
5
15
15
32
480
43.79±0.29
RRM
5
15
21
32
672
43.79±0.14
RRM
0
15
15
2
30
45.79±0.29
Appendix
Table 8: ARC-AGI-1 ablations. Tdetach is the number of initial applications, out of Ttrain , through which gradients are truncated. Total gives Ttest×Nsup . All models use L=2 . The TRM-Att result is from Jolicoeur-Martineau (2025) .
Recursive reasoning models apply a small shared Transformer block many times to refine a latent state. This gives them large effective depth with few parameters and makes them strong on algorithmic tasks. Such compact solvers are natural candidates for tools that an LLM can call on narrow algorithmic subproblems. However, existing models such as HRM, TRM and URM differ in architecture, gradient propagation and training procedure simultaneously. This makes it hard to tell what drives their performance, and their optimization is still poorly understood and often unstable. In this work we address both of these gaps. First, we study these questions under a unified experimental pipeline spanning six algorithmic domains. Individual controlled ablations are performed on representative domains, while the resulting recipe is evaluated across the full suite. The study reveals a surprisingly simple recipe for stable and generalizable recursive reasoning: an intermediate gradient horizon, large physical batches and controlled updates of the recurrent state. An explicit hierarchical architecture is not needed. Second, we combine these findings into a stable 13.6M-parameter model that achieves the strongest overall performance among the evaluated recursive baselines, with particularly large gains on out-of-distribution generalization. It raises Arithmetic OOD accuracy to 71.2%, from 36.2% for the strongest baseline, while reaching 98.41% on Sudoku and 59.5% pass@2 on ARC-AGI-1. Our results show that, within the recursive architectures studied here, performance depends strongly on how recurrence is optimized and stabilized. More broadly, it shows how AI systems can be improved by optimizing their components one at a time.
Tiny Recursive Models (TRM) solve complex reasoning tasks with a fraction of the parameters of modern large language models (LLMs) by iteratively refining a latent state and final answer. While powerful, their deterministic recursion can lead to convergence at suboptimal solutions, without escape mechanism. A common workaround relies on task-specific input perturbations at test time combined with answer aggregation via voting. We introduce Probabilistic TRM (PTRM), a task-agnostic framework for test-time compute scaling that addresses this limitation through stochastic exploration. PTRM injects Gaussian noise at each deep recursion step, enabling parallel trajectories to explore diverse solution basins, and selects among them using the model's existing Q head (used for early stopping in the original TRM). Without requiring retraining or task-specific augmentations, PTRM enables substantial accuracy gains across benchmarks, including Sudoku-Extreme (87.4% to 98.75%) and on various puzzles from Pencil Puzzle Bench (62.6% to 91.2%). On the latter, PTRM achieves nearly double the accuracy of frontier LLMs (91.2% vs. 55.1%) at less than 0.0001x the cost, using only 7M parameters.
Amin Sghaier, Ali Parviz, Alexia Jolicoeur-Martineau
Mila – Quebec AI Institute ILLS & ETS Montreal · Independent
How should future neural reasoning systems implement extended computation? Recursive Reasoning Models (RRMs) offer a promising alternative to autoregressive sequence extension by performing iterative latent-state refinement with shared transition functions. Yet existing RRMs are largely deterministic, following a single latent trajectory and converging to a single prediction. We introduce Generative Recursive reAsoning Models (GRAM), a framework that turns recursive latent reasoning into probabilistic multi-trajectory computation. GRAM models reasoning as a stochastic latent trajectory, enabling multiple hypotheses, alternative solution strategies, and inference-time scaling through both recursive depth and parallel trajectory sampling. This yields a latent-variable generative model supporting conditional reasoning via pθ(y∣x) and, with fixed or absent inputs, unconditional generation via pθ(x). Trained with amortized variational inference, GRAM improves over deterministic recurrent and recursive baselines on structured reasoning and multi-solution constraint satisfaction tasks, while demonstrating an unconditional generation capability. https://ahn-ml.github.io/gram-website
Junyeob Baek, Mingyu Jo, Minsu Kim +3
KAIST · Mila – Québec AI Institute · New York University +1