cs.AISep 27, 2026

Not Too Hard, Not Too Easy: Learning from Intermediate States for LLM Structured Reasoning

Authors: Hongbo Chen, Guohua Lu, Ting Dang, Hong Jia

Organizations: University of Auckland, Auckland, New Zealand · University of Melbourne, Melbourne, Australia

Abstract

A common principle of effective learning is to practice material that is neither already mastered nor too difficult to permit progress. We ask how to apply this principle to structured reasoning tasks such as Sudoku and maze solving. In these tasks, a model can repeatedly revise an incomplete or incorrect candidate solution until it satisfies the problem's constraints. The intermediate candidate solutions along this trajectory provide natural training examples: some are already solved, some cannot yet be repaired by the model, and others lie at its current frontier of achievable progress. We therefore investigate whether pretrained language models can learn to revise such states and whether training on states at this frontier improves reasoning more broadly. To achieve this, we couple a pretrained language-model backbone with a recurrent updater that repeatedly revises an explicit solution state, using the same parameters at every update step. We further introduce Frontier-Oriented Curation Using Self-trajectories (FOCUS), which selects training states from trajectories generated by the current model. FOCUS measures how much the model improves each state within a fixed number of recurrent updates and prioritizes states from which it can make substantial progress. With Qwen3-1.7B, FOCUS achieves 64.4% exact solve accuracy on Sudoku-Extreme and 91.1% on Maze-Hard, with similar gains observed across five Qwen and Llama backbones spanning 1.7B to 8B parameters. We further observe zero-shot transfer in the adapted LLM to mathematical reasoning and code execution, even when the recurrent updater is disabled and no downstream fine-tuning is performed.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Frontier Learning: Training LLM Reasoners at the Edge of Capability

    Sep 28, 2026Robin Faro, Shyam Sundhar Ramesh, Ilija Bogunovic +1LLM Reasoning StrategiesPost-Training

  2. Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces

    Sep 22, 2026Minghui Liu, Thomas Magelinski, Dehao Yuan +2LLM Reasoning StrategiesThoughts

  3. Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning

    Sep 22, 2026Ismail Labiad, Matthieu Kowalski, Marc Schoenauer +2LLM Reasoning StrategiesReasoning Skills