cs.AIOct 8, 2026

Learning to Plan by Looking Back: Hindsight Hierarchies for Training Reasoning Models

Authors: Lars Simon, Holger Eble, Manuel Radons

Organizations: Bundesdruckerei GmbH Berlin, Germany

Abstract

We introduce a self-improvement loop for reasoning models based on the following observation: Even when the difficulty of a problem exceeds the model's current solving abilities, an additionally supplied solution might enable the model to extract useful solution ideas in hindsight. We operationalize this by jointly training the same model to exhibit the following three capabilities: predicting solution ideas from problems alone, reverse-engineering ideas from problems and known solutions, and solving problems using provided ideas. The loop alternates between reverse engineering such ideas from problems with supplied solutions and using these ideas as additional supervision for joint training of all three capabilities. We give a formal specification of our method and a concrete instantiation for interactive theorem proving in the Lean theorem prover; empirical evaluation remains future work.

Figures & tables

Explore similar work

CardsList
  1. ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling

    Aug 11, 2026Vaibhav Singh, Soumya Suvra Ghosal, Sarvesh Gharat +3Numerical Reasoning in Language ModelsTest-Time Scaling

  2. Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning

    Sep 22, 2026Ismail Labiad, Matthieu Kowalski, Marc Schoenauer +2RL for Language Model ReasoningLLM Reasoning

  3. Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning

    Jul 8, 2026Dmitry Beresnev, Vladimir Makharev, Roman Khalikov +2LLM Self-CorrectionInference-Time Search