We introduce a self-improvement loop for reasoning models based on the following observation: Even when the difficulty of a problem exceeds the model's current solving abilities, an additionally supplied solution might enable the model to extract useful solution ideas in hindsight. We operationalize this by jointly training the same model to exhibit the following three capabilities: predicting solution ideas from problems alone, reverse-engineering ideas from problems and known solutions, and solving problems using provided ideas. The loop alternates between reverse engineering such ideas from problems with supplied solutions and using these ideas as additional supervision for joint training of all three capabilities. We give a formal specification of our method and a concrete instantiation for interactive theorem proving in the Lean theorem prover; empirical evaluation remains future work.
Figures & tables
Figure 1: Simplified schematic of the training loop. Each iteration generates hindsight ideas for a new batch and jointly trains the same model to predict, reconstruct, and use ideas.
Figure 2: Hierarchy generation in both directions (two levels, n=1 ) and inference. The solver receives the full predicted hierarchy. The example texts are hand-written illustrations, not model outputs or empirical observations.
Figure 3: Schematic for the full training loop and the two choices for ConstructTriple from Sections 4.5 and 4.6 . The assessment box illustrates one possible way to score candidate hierarchies and retain a valid solution.
Figure 4: Hierarchy-guided Lean inference with J sequential arms. Each arm samples a hierarchy from Fθ and subsequently invokes the solver with this fixed hierarchy and a fresh proof search graph rooted at the initial state s0 . An unsuccessful arm is followed by the next and a valid solution stops the entire procedure; we signal failure if all J arms have been exhausted. The right panel illustrates tactic generation and Lean interaction within one arm.