Latent reasoning models repeatedly update a latent state using the same recurrent block. As the recurrent depth increases, the sequence of latent states may converge to a compact subset of the state space without necessarily converging to a fixed point. Existing models typically predict by applying a prediction head to a single latent state. However, the latent state can continue to change even after many updates, potentially making predictions unstable across recurrent depths. To address this instability, we introduce invariant-measure reasoners (ImR), a framework that uses an invariant measure as a stable representation. This measure describes the long-run distribution of latent states on the compact subset and is invariant under updates by the recurrent block. ImR predicts from the expectation of the prediction head's output under this measure. We use ImR in two ways: fine-tuning only the prediction head of existing models and training models from scratch. Both approaches reduce prediction instability and improve accuracy in many settings on maze and Sudoku tasks. In some settings, models trained with ImR exhibit non-fixed-point behavior more frequently than existing models yet achieve high accuracy even with such behavior, unlike existing models. These results suggest that ImR can leverage otherwise destabilizing dynamics for latent reasoning.
Figures & tables
Figure 1: (a) An example of prediction instability in equilibrium reasoners (EqR) ( Huang et al., 2026 ) on a maze problem with the latent trajectory projected onto the first two principal components. Trajectory color indicates recurrent depth, and pale green and red regions approximately indicate correct and incorrect predictions, respectively. (b) Schematic of ImR. For an input x , repeated applications of the recurrent block fx to the latent state st generate a latent trajectory. ImR predicts from the expectation Es∼μ⋆[g(s)] under an invariant measure μ⋆ .
Figure 2: Accuracy across recurrent depth on Maze-Unique (a) and Sudoku-Extreme (b), with enlarged views (c) over the final 25 recurrent depths. For both fine-tuned models and ImR, curves show means and shaded regions show standard deviation across 3 training seeds.
Fixed
Non-fixed
Model
Non-fixed rate
Stable correct
Stable incorrect
Unstable
Stable correct
Stable incorrect
Unstable
Sudoku-Extreme : dimension of s(H) : 49.7K ; parameters 5.03M
EqR †
4.1
100.0
0.0
0.0
3.6
0.0
96.4
ImR (FT)
4.1±0.0
99.2±0.0
0.0±0.0
0.8±0.0
3.4±0.0
79.7±0.1
16.9±0.1
ImR
1.2±0.2
99.7±0.1
0.0±0.0
0.2±0.1
8.0±12.2
44.5±5.8
47.5±8.0
Sudoku-Extreme : dimension of s(H) : 41.5K ; parameters 4.86M
Table 1: Trajectory groups and prediction stability. Entries with ± report mean and standard deviation across 3 training seeds. EqR † reports results for both EqR and EqR (FT), which match on all reported metrics. For each dataset and state dimension of s(H) , bold marks the highest stable correct proportion and the lowest unstable proportion in the non-fixed group.
Model
Stable correct (%)
Stable incorrect (%)
Unstable (%)
Degree of the start cell: 2
DT
76.6
17.3
6.1
DT (FT)
81.5±0.1
13.9±0.1
4.5±0.1
ImR (FT)
84.5±0.0
15.1±0.1
0.4±0.1
Degree of the start cell: 3
DT
25.3
57.9
16.8
Table 2: Prediction stability on OOD mazes over depths 4,900 – 5,000 . Fine-tuned models report mean ± standard deviation across 3 training seeds. For each degree of the start cell, bold marks the highest stable correct proportion and the lowest unstable proportion.
Figure 3: Accuracy across recurrent depth on OOD mazes whose start cell has degree 2 (a) or 3 (b), with enlarged views (c) over the final 20 recurrent depths. For both fine-tuned models, curves show means and shaded regions show standard deviation across 3 training seeds.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Train
Validation
Test
Maze-Unique
1,000
1,000
1,000
Sudoku-Extreme
1,001,000
10,000
422,786
Maze
40,000
10,000
1,000
Appendix
Table 3: Dataset sizes (top) and counts of generated test mazes by the degree of the start cell (bottom). The degree counts are pooled across all maze sizes. The Sudoku-Extreme training count includes augmentation.
Figure 4: Examples of ID and OOD mazes used for DT. The first column shows a 9×9 ID validation maze. The remaining columns show 19×19 OOD mazes whose start cells have degrees 1 through 4. The top row contains the inputs, and the bottom row marks the target solution path in blue, the start cell in green, and the goal cell in red.
Dataset
Position mixing
Puzzle prefix
fx
g
Total
Maze-Unique
Self-attention
With
263,040
768
264,066
Without
262,912
768
263,938
Sudoku-Extreme
MLP
With
5,022,720
5,632
5,029,378
Without
4,848,640
5,632
4,855,298
Maze
Convolution
Without
740,736
39,312
783,504
Appendix
Table 4: Model architectures and the number of parameters. For EqR, parameters used to embed the input are included under fx . The totals also include the halting head in EqR and the neural network used to initialize the latent state in DT.
Architecture
Dataset
Batch size
Optimizer updates
EqR
Maze-Unique
768
1,000
EqR
Sudoku-Extreme
768
1,000
DT
Maze
50
1,600
Appendix
Table 5: Hyperparameters for fine-tuning the prediction head.
Model
Dataset
Puzzle prefix
L
K
Batch size
Optimizer updates
ImR
Maze-Unique
With / Without
36
4
768
150,000
ImR
Sudoku-Extreme
With / Without
36
4
768
50,000
EqR
Maze-Unique
Without
36
4
768
150,000
EqR
Sudoku-Extreme
Without
3
1
768
50,000
Appendix
Table 6: Hyperparameters selected for training from scratch with the architecture of EqR. Here, L is the number of recurrent updates per inner loop, and gradients through the recurrent block are retained for the final K updates.
Figure 5: The EqR trajectory in Figure (a) and its actual predictions. Left: the trajectory over depths 0–512 and prediction regions on the reconstructed PCA plane. Trajectory color indicates recurrent depth; the circle marks the initial state. Right: whole-maze correctness from the original head applied to each actual latent state. Pale green and pale red indicate correct and incorrect predictions, respectively.
Model
Puzzle prefix
Accuracy (%)
Maze-Unique , T=512
EqR
With
89.4
EqR (FT)
With
89.4±0.0
ImR (FT)
With
91.6±0.1
ImR
With
98.8±0.9
EqR
Without
82.9±20.9
Appendix
Table 7: Accuracy at the final recurrent depth. Entries with ± report the mean and standard deviation across 3 training seeds. The remaining entries use a single released model.
Figure 6: PCA projections of the part of the latent state used by the prediction head for ImR models trained without the puzzle prefix. Each row shows an input and the corresponding latent trajectory projected onto PC1–PC2, PC3–PC4, and PC5–PC6. The first two rows show Maze-Unique examples, and the last two show Sudoku-Extreme examples. Axis labels give the percentage of variance explained by each principal component. Colors indicate recurrent depths 897–1024; circles and stars mark the first and last states in this interval, respectively. All trajectories were generated with noise set to zero.
Figure 7: Accuracy across recurrent depth for each maze size. The first two columns show sizes 9×9 through 99×99 , with 100 test mazes per size, combining all degrees of the start cell. The right column shows enlarged views over depths 4,981 – 5,000 . Curves for DT (FT) and ImR (FT) report means across 3 training seeds.