Looped reasoners spend test-time compute by iterating a weight-tied map, but a small residual does not mean the state is a fixed point when that map lives in unconstrained latent space. We propose Geometric Fixed-Point Reasoning (GFPR), in which the iterated state is the prediction itself: a field of categorical beliefs on a product of simplices, whose argmax is the answer at every step. Because the state is a belief, task structure can be imposed through compact convex relaxations, either as structured readouts or directly in the recurrent state; in the latter case the update remains a continuous self-map, so a fixed point exists for any parameters. At about 7M parameters, GFPR reaches 95.1% exact match on Sudoku-Extreme, 92.0% on Maze-Hard, and 100% sequence accuracy on S_5 length 128, above the published FPRM numbers at the same scale. The same update also trains a 201M language model on FineWeb-Edu in which each site is a distribution over the vocabulary; with 24 Picard steps it is above GPT-2 small on four zero-shot multiple-choice tasks and above GPT-2 medium on ARC-Easy.
Figures & tables
Figure 1: In GFPR the state is a field of simplex-valued beliefs: at each site, a distribution over local outputs (e.g. Sudoku digits or vocabulary tokens) together with auxiliary register coordinates that store internal state but are not decoded as answers. At test time we unroll damped Picard steps; the prediction is argmax over the output block of the final belief, with known outputs pinned throughout.
Figure 2: One randomly chosen Sudoku-Extreme test puzzle that both FPRM and GFPR solve exactly. The dotted line marks where FPRM’s own rule halts ( k=36 ); in the shaded region we keep running FPRM to see whether it settles. (a) One-step residual ∥F(u)−u∥/∥u∥ in float64. GFPR’s residual falls to float64 precision; FPRM’s stays near 10−2 before and after its halt. (b) Local amplification probe at each step: Kpi=20 power iterations on JVPs (float32), plotted as ∥JFv∥ after renormalization (not a certified ρ(JF) when JF is non-normal; section G ). The dashed line marks unit amplification. On the final states of this puzzle only, restarted Arnoldi gives largest Ritz moduli ≈1.3×10−3 (GFPR) and ≈1.04 (FPRM).
Model
Params
Sudoku-Extreme
Maze-Hard
S5
HRM
27M
55.0
74.5
–
TRM (attention)
7M
74.7
85.3
39.4
TRM (MLP)
5M
87.4
–
–
TRM + causal conv
7M
–
–
97.2
EqR
7M
93.0 †
–
–
FPRM
7M
94.2
87.0
98.8
Table 1: Exact-match accuracy (%). “–” was not reported. Quoted baselines are from Movahedi et al. (2026) , except TRM-MLP ( Jolicoeur-Martineau, 2025 ) . † One trajectory. ‡ A sample of 23,680 Sudoku test puzzles. Slash-separated parameter counts correspond to Sudoku-Extreme, Maze-Hard, and S5 , in that order.
GFPR-LM, Picard steps K
Task
GPT-2 S
GPT-2 M
1
8
16
24
ARC-Easy
44.2
49.2
32.0
28.4
49.5
51.9
SciQ
72.5
75.3
37.4
32.9
76.1
74.4
PIQA
62.4
66.9
55.6
53.8
62.8
63.1
HellaSwag
31.2
39.5
27.8
26.0
32.2
33.2
Table 2: Zero-shot multiple-choice accuracy (%). ARC-Easy, SciQ, and PIQA are accuracy; HellaSwag is length-normalized accuracy. GFPR-LM is the 201M FineWeb-Edu run (EMA weights, β=0.5 , 6B tokens), scored by next-token log-likelihood after K damped Picard steps from the uniform state. GPT-2 small (124M) and medium (355M) use one forward pass in the same harness. All GFPR-LM columns are official full-split figures for this checkpoint (sweep through K=32 in section F ); K=24 is the training depth, not a test-set choice. Bold marks the best of GPT-2 and K=24 ; the K=8 and K=16 columns show how the scores depend on depth.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: One test-time step. Arrows are the loop. The bottom row is a property of that loop, not another update. FPRM reads the grid from a head on an arbitrary latent. GFPR reads the grid from the belief it iterates, after pinning the given clues.
Sudoku-Extreme
Maze-Hard
S5
Parameters
7.16M
6.81M
5.60M
Width / layers / heads
256 / 9 / 8
512 / 2 / 8
224 / 9 / 8
K / A
9 / 16
6 / 8
121 / 16
Self-conditioning passes
4
2
4
Damping β
0.7
0.3
0.7
Rollout depth
Dˉ=32
40
Dˉ=32
Appendix
Table 3: GFPR settings for the released runs. K is the number of output symbols and A the number of auxiliary coordinates; for S5, K=121 counts the 120 permutations and the pad token. Rollout depth is the no-gradient prefix: “ Dˉ=32 ” is the log-normal Poisson draw in the main text, and the TV tolerance is its early-stop threshold. “cosine@150k” starts the learning-rate decay at step 150,000. S5 applies a causal depthwise convolution of kernel size 4 to the state before adding operation embeddings; the other runs do not.
Figure 7
Task
K=1
2
4
8
16
24
32
ARC-Easy
32.0
32.1
29.4
28.4
49.5
51.9
51.9
SciQ
37.4
34.6
29.9
32.9
76.1
74.4
74.3
PIQA
55.6
55.6
54.5
53.8
62.8
63.1
63.2
HellaSwag
27.8
27.6
26.8
26.0
32.2
33.2
33.1
Appendix
Table 5: GFPR-LM 201M zero-shot accuracy (%) versus Picard depth K . HellaSwag is length-normalized accuracy.