A common principle of effective learning is to practice material that is neither already mastered nor too difficult to permit progress. We ask how to apply this principle to structured reasoning tasks such as Sudoku and maze solving. In these tasks, a model can repeatedly revise an incomplete or incorrect candidate solution until it satisfies the problem's constraints. The intermediate candidate solutions along this trajectory provide natural training examples: some are already solved, some cannot yet be repaired by the model, and others lie at its current frontier of achievable progress. We therefore investigate whether pretrained language models can learn to revise such states and whether training on states at this frontier improves reasoning more broadly. To achieve this, we couple a pretrained language-model backbone with a recurrent updater that repeatedly revises an explicit solution state, using the same parameters at every update step. We further introduce Frontier-Oriented Curation Using Self-trajectories (FOCUS), which selects training states from trajectories generated by the current model. FOCUS measures how much the model improves each state within a fixed number of recurrent updates and prioritizes states from which it can make substantial progress. With Qwen3-1.7B, FOCUS achieves 64.4% exact solve accuracy on Sudoku-Extreme and 91.1% on Maze-Hard, with similar gains observed across five Qwen and Llama backbones spanning 1.7B to 8B parameters. We further observe zero-shot transfer in the adapted LLM to mathematical reasoning and code execution, even when the recurrent updater is disabled and no downstream fine-tuning is performed.
Figures & tables
Figure 1: Overview of FOCUS. (a) Training selects intermediate states near the repair frontier and resumes recurrent updates from the selected answer and memory. (b) At inference, the adapted LLM encodes the input once to condition K recurrent updates, followed by a single final decode.
Backbone
Sudoku-Extreme
Maze-Hard
Direct
One-step
Recurrent
FOCUS
Direct
One-step
Recurrent
FOCUS
Qwen3-1.7B
0.5
0.5
7.2
64.4
0.0
0.0
86.0
91.1
Qwen3-4B
0.5
0.3
16.2
66.0
0.1
0.0
86.9
91.5
Qwen3-8B
0.6
0.7
14.8
64.7
0.0
0.0
87.2
88.3
Llama-3.2-3B-Instruct
0.0
0.4
12.3
61.9
0.0
0.0
86.8
90.1
Llama-3.1-8B-Instruct
0.3
0.8
8.0
65.8
0.0
0.0
79.8
87.7
Table 1: Source-task accuracy across language-model backbones. Exact solve accuracy (%) on Sudoku-Extreme and Maze-Hard.
Table 3: Zero-shot transfer to AIME25 and AIME26 after structured-repair training.
Mathematics
Code execution
Source-task adaptation
MATH
MATH- Hard
CruxEval- O
CruxEval- I
Avg.
Qwen3-1.7B
Baseline
86.54
72.36
50.75
24.00
58.41
FOCUS (Sudoku)
86.38
72.96
55.38
26.75
60.37
FOCUS (Maze)
86.36
71.68
50.63
24.12
58.20
Qwen3-4B
Table 4: Zero-shot transfer after structured-repair training. Accuracy (%) with one greedy completion per item. The recurrent module is disabled and no downstream parameters are updated.
Method
Replay-state selection
Sudoku- Extreme
Maze- Hard
Random
Uniformly sampled rollout state
57.8
89.7
Energy-Hard
Rollout state with the highest task error
61.3
89.9
Fixed Mix
Fixed mixture of training-state types
60.8
90.5
FOCUS
Rollout state nearest the repair frontier
64.4
91.1
Table 5: State-curation comparison with Qwen3-1.7B. Held-out exact solve accuracy (%) on Sudoku-Extreme ( K=128 ) and Maze-Hard ( K=16 ).
Method
Sudoku-Extreme
Maze-Hard
FOCUS w/o defect penalty
61.3
85.8
FOCUS
64.4
91.1
Table 6: Defect-penalty ablation. Qwen3-1.7B exact solve accuracy (%) with and without the penalty.
Figure 2: Recurrent repair with Qwen3-1.7B FOCUS. (a) The Maze-Hard example first matches the target at step 16. (b) The Sudoku-Extreme example first matches the target at step 34. Maze path counts exclude the start and goal; Sudoku cell accuracy includes all 81 cells.
Figure 3: Maze-Hard predictions with Qwen3-8B. Each row starts with the input. Direct and One-step show final predictions; Recurrent final-only and FOCUS show steps 4, 7, and 16. The recurrent baseline ends with a valid but non-shortest path; FOCUS matches the reference at step 7 and retains it at step 16. Insets A and B magnify the boxed regions at step 16. Colors indicate agreement with the reference after inference.
0.80 (multiplier on the predicted logit increment)
0.50 (multiplier on the predicted logit increment)
Appendix
Table 7: Training settings for Qwen3-1.7B FOCUS on Sudoku-Extreme and Maze-Hard.
Method
Sudoku-Extreme
Maze-Hard
FOCUS with random selection
61.1
87.1
FOCUS
64.4
91.1
Appendix
Table 8: State-selection ablation. Qwen3-1.7B exact solve accuracy (%) on Sudoku-Extreme ( K=128 ) and Maze-Hard ( K=16 ).
Figure 4: Visualization of a Sudoku puzzle and its solving steps using FOCUS.
Figure 5: Visualization of a maze and its solving steps using FOCUS.
Figure 6: Predictions on the same Sudoku-Extreme test puzzle using Qwen3-8B. Each row starts with the input. Direct LoRA SFT and One-step State-SFT show their final answers in the second position; Recurrent final-only and FOCUS show predictions at steps 4, 8, and 128. FOCUS solves this example at step 8 and retains the solution through step 128. Insets A and B magnify the boxed regions at step 128. Colors indicate agreement with the target after inference.
Figure 7: Baseline and FOCUS response excerpts on MATH (example 1).
Figure 8: Baseline and FOCUS response excerpts on MATH (example 2).
Figure 9: Baseline and FOCUS response excerpts on AIME 2026 (example 1).
Figure 10: Baseline and FOCUS response excerpts on AIME 2026 (example 2).
Figure 11: Baseline and FOCUS response excerpts on MATH-Hard (example 1).
Figure 12: Baseline and FOCUS response excerpts on MATH-Hard (example 2).
Figure 13: Baseline and FOCUS response excerpts on CruxEval-O (example 1).
Figure 14: Baseline and FOCUS response excerpts on CruxEval-O (example 2).
Figure 15: Baseline and FOCUS response excerpts on CruxEval-I (example 1).
Figure 16: Baseline and FOCUS response excerpts on CruxEval-I (example 2).