A common principle of effective learning is to practice material that is neither already mastered nor too difficult to permit progress. We ask how to apply this principle to structured reasoning tasks such as Sudoku and maze solving. In these tasks, a model can repeatedly revise an incomplete or incorrect candidate solution until it satisfies the problem's constraints. The intermediate candidate solutions along this trajectory provide natural training examples: some are already solved, some cannot yet be repaired by the model, and others lie at its current frontier of achievable progress. We therefore investigate whether pretrained language models can learn to revise such states and whether training on states at this frontier improves reasoning more broadly. To achieve this, we couple a pretrained language-model backbone with a recurrent updater that repeatedly revises an explicit solution state, using the same parameters at every update step. We further introduce Frontier-Oriented Curation Using Self-trajectories (FOCUS), which selects training states from trajectories generated by the current model. FOCUS measures how much the model improves each state within a fixed number of recurrent updates and prioritizes states from which it can make substantial progress. With Qwen3-1.7B, FOCUS achieves 64.4% exact solve accuracy on Sudoku-Extreme and 91.1% on Maze-Hard, with similar gains observed across five Qwen and Llama backbones spanning 1.7B to 8B parameters. We further observe zero-shot transfer in the adapted LLM to mathematical reasoning and code execution, even when the recurrent updater is disabled and no downstream fine-tuning is performed.
Figures & tables
Figure 1: Overview of FOCUS. (a) Training selects intermediate states near the repair frontier and resumes recurrent updates from the selected answer and memory. (b) At inference, the adapted LLM encodes the input once to condition K recurrent updates, followed by a single final decode.
Backbone
Sudoku-Extreme
Maze-Hard
Direct
One-step
Recurrent
FOCUS
Direct
One-step
Recurrent
FOCUS
Qwen3-1.7B
0.5
0.5
7.2
64.4
0.0
0.0
86.0
91.1
Qwen3-4B
0.5
0.3
16.2
66.0
0.1
0.0
86.9
91.5
Qwen3-8B
0.6
0.7
14.8
64.7
0.0
0.0
87.2
88.3
Llama-3.2-3B-Instruct
0.0
0.4
12.3
61.9
0.0
0.0
86.8
90.1
Llama-3.1-8B-Instruct
0.3
0.8
8.0
65.8
0.0
0.0
79.8
87.7
Table 1: Source-task accuracy across language-model backbones. Exact solve accuracy (%) on Sudoku-Extreme and Maze-Hard.
Table 3: Zero-shot transfer to AIME25 and AIME26 after structured-repair training.
Mathematics
Code execution
Source-task adaptation
MATH
MATH- Hard
CruxEval- O
CruxEval- I
Avg.
Qwen3-1.7B
Baseline
86.54
72.36
50.75
24.00
58.41
FOCUS (Sudoku)
86.38
72.96
55.38
26.75
60.37
FOCUS (Maze)
86.36
71.68
50.63
24.12
58.20
Qwen3-4B
Table 4: Zero-shot transfer after structured-repair training. Accuracy (%) with one greedy completion per item. The recurrent module is disabled and no downstream parameters are updated.
Method
Replay-state selection
Sudoku- Extreme
Maze- Hard
Random
Uniformly sampled rollout state
57.8
89.7
Energy-Hard
Rollout state with the highest task error
61.3
89.9
Fixed Mix
Fixed mixture of training-state types
60.8
90.5
FOCUS
Rollout state nearest the repair frontier
64.4
91.1
Table 5: State-curation comparison with Qwen3-1.7B. Held-out exact solve accuracy (%) on Sudoku-Extreme ( K=128 ) and Maze-Hard ( K=16 ).
Method
Sudoku-Extreme
Maze-Hard
FOCUS w/o defect penalty
61.3
85.8
FOCUS
64.4
91.1
Table 6: Defect-penalty ablation. Qwen3-1.7B exact solve accuracy (%) with and without the penalty.
Figure 2: Recurrent repair with Qwen3-1.7B FOCUS. (a) The Maze-Hard example first matches the target at step 16. (b) The Sudoku-Extreme example first matches the target at step 34. Maze path counts exclude the start and goal; Sudoku cell accuracy includes all 81 cells.
Figure 3: Maze-Hard predictions with Qwen3-8B. Each row starts with the input. Direct and One-step show final predictions; Recurrent final-only and FOCUS show steps 4, 7, and 16. The recurrent baseline ends with a valid but non-shortest path; FOCUS matches the reference at step 7 and retains it at step 16. Insets A and B magnify the boxed regions at step 16. Colors indicate agreement with the reference after inference.
0.80 (multiplier on the predicted logit increment)
0.50 (multiplier on the predicted logit increment)
Appendix
Table 7: Training settings for Qwen3-1.7B FOCUS on Sudoku-Extreme and Maze-Hard.
Method
Sudoku-Extreme
Maze-Hard
FOCUS with random selection
61.1
87.1
FOCUS
64.4
91.1
Appendix
Table 8: State-selection ablation. Qwen3-1.7B exact solve accuracy (%) on Sudoku-Extreme ( K=128 ) and Maze-Hard ( K=16 ).
Figure 4: Visualization of a Sudoku puzzle and its solving steps using FOCUS.
Figure 5: Visualization of a maze and its solving steps using FOCUS.
Figure 6: Predictions on the same Sudoku-Extreme test puzzle using Qwen3-8B. Each row starts with the input. Direct LoRA SFT and One-step State-SFT show their final answers in the second position; Recurrent final-only and FOCUS show predictions at steps 4, 8, and 128. FOCUS solves this example at step 8 and retains the solution through step 128. Insets A and B magnify the boxed regions at step 128. Colors indicate agreement with the target after inference.
Figure 7: Baseline and FOCUS response excerpts on MATH (example 1).
Figure 8: Baseline and FOCUS response excerpts on MATH (example 2).
Figure 9: Baseline and FOCUS response excerpts on AIME 2026 (example 1).
Figure 10: Baseline and FOCUS response excerpts on AIME 2026 (example 2).
Figure 11: Baseline and FOCUS response excerpts on MATH-Hard (example 1).
Figure 12: Baseline and FOCUS response excerpts on MATH-Hard (example 2).
Figure 13: Baseline and FOCUS response excerpts on CruxEval-O (example 1).
Figure 14: Baseline and FOCUS response excerpts on CruxEval-O (example 2).
Figure 15: Baseline and FOCUS response excerpts on CruxEval-I (example 1).
Figure 16: Baseline and FOCUS response excerpts on CruxEval-I (example 2).
Reinforcement Learning-based post-training of Large Language Models (LLM) has been successfully applied to improve their reasoning capabilities. Existing pipelines primarily finetune LLMs on a fixed pool of problems specified prior to training using the GRPO loss. This is fundamentally limiting, as learning signal arises only when policy rollouts mix successes and failures, causing the useful portion of any fixed pool to quickly become stale as the model improves. To address this, we propose frontier learning, an open-ended post-training approach in which procedural generators are used online to continually produce informative training problems. It treats the generator's task-specific parameters as a search space and uses a regret signal to prioritize and explore frontier difficulty levels in order to focus training at the edge of the model's evolving reasoning capabilities. Across several reasoning tasks and model families, our approach consistently achieves higher relative gains over fixed-pool baselines, demonstrating that effective post-training requires not only selecting useful problems, but continually generating them at the edge of capability.
Robin Faro, Shyam Sundhar Ramesh, Ilija Bogunovic +1
University College London London, UK · University of Basel Basel, Switzerland
Large language models (LLMs) excel at reasoning when scaled to hundreds of billions of parameters, but small- and mid-scale models remain brittle reasoners even with knowledge distillation (KD). We present Ladders-of-Thought (LoT), a framework that improves reasoning by combining progressive question rewrites with a self-evolving curriculum. LoT automatically generates semantically faithful but easier variants of reasoning problems, organizes them into difficulty buckets using step-based measures, and employs a self-evolving bandit scheduler to allocate training adaptively. Evaluated on two reasoning domains, math and multi-hop reasoning, across 1-8B models from different families, LoT consistently improves over KD. It delivers large gains on arithmetic tasks (e.g., +32 percentage points on AddSub, +25pp on SVAMP), +2-8pp improvements on in-domain test splits, and strong though dataset-dependent benefits on multi-hop reasoning (e.g., +16pp on QASC, +25pp on StrategyQA). LoT also converges faster than staged curricula, highlighting the value of adaptive progression. These results show that progressive rewrites coupled with adaptive curricula provide a simple yet effective recipe for strengthening reasoning in smaller LLMs.
Minghui Liu, Thomas Magelinski, Dehao Yuan +2
University of Maryland, College Park · Capital One
Large language models increasingly tackle hard reasoning problems by spending more test-time compute, yet the dominant strategy remains naive repeated sampling: draw many independent solutions and hope one is correct. Because such sampling explores only through local decoding noise, it tends to produce many near duplicate attempts rather than genuinely different ideas. We ask whether exploration can instead be steered at a semantic level, by first sampling problem specific concepts, hints, or strategies and then conditioning answer generation on them. We refine this into a simple, more exploratory procedure that emits many diverse concepts in a single trajectory, and evaluate it on hard problems where repeated sampling struggles. We then go a step further and make concept generation trainable: a small concept generator is optimized with reinforcement learning so that its concepts maximize the downstream success of a larger, frozen answer generator. On hard mathematical reasoning problems, the trained concept generator substantially improves the answer generator's pass@k over naive repeated sampling at the same answer generation allocation, surpasses concepts drawn from much larger untuned models, and transfers to answer generators it was never trained against, including a model from a different family. A small model can thus be trained into an effective, reusable search policy for a much larger one.
Ismail Labiad, Matthieu Kowalski, Marc Schoenauer +2
Meta FAIR · Université Paris-Saclay, LISN, Inria, CNRS · NYU Courant Institute and CDS