cs.AIOct 8, 2026

Balancing Reference Guidance and Free Generation in Trajectory Rollouts for Reasoning RL

Authors: Hanyu Wang, Nakul Agarwal, Hossein Nourkhiz Mahjoub, Ehsan Moradi Pari, Makoto Fukushima, Jinghui Chen, Vaishnav Tadiparthi

Organizations: The Pennsylvania State University · Honda Research Institute USA · Honda Research Institute Japan

Abstract

A verified reference solution provides a correct trajectory for training a reasoning model. Alternatively, a prefix of the reference can guide the model in generating a trajectory of its own. How much reference guidance should we provide? We study this question through prefix continuation, where the model continues from a reference prefix and keeps the resulting trajectory if it passes verification, falling back to the reference otherwise. Since both procedures produce correct trajectories, we compare their distributions with the ideal distribution, the model's own distribution conditioned on successful verification. For one continuation, we derive the KL divergence in closed form, which, up to a bounded term, decreases with the product of the probability of generating a different correct trajectory and the reference surprisal, the negative log probability of the reference suffix given the prefix. Since a longer prefix tends to raise the former but lowers the latter, continuation success alone does not determine the preferred amount of guidance. From this analysis, we learn a prefix selector shared across training questions from continuation outcomes, without estimating success probabilities or additional generation. The resulting Adaptive Reference Guidance (ARG) constructs correct trajectories within a fixed generation budget, and we apply it to all-failure groups in Group Relative Policy Optimization (GRPO). Experiments on Qwen3-4B and Qwen3-8B across five mathematical reasoning benchmarks show that ARG achieves the highest aggregate pass@12 among the evaluated methods with competitive average sampled accuracy.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

    Jul 21, 2026Priyank Agrawal, Ankur Samanta, Shervin Ghasemlou +4Reinforcement LearningRL for Language Model Reasoning

  2. ExTra: Exploratory Trajectory Optimization for Language Model Reinforcement Learning

    Jun 23, 2026Wenyang Hu, Junxiang Jia, Zhen Shu +3RL for Language ModelsRL for Language Model Reasoning

  3. Selective Off-Policy Reference Tuning with Plan Guidance

    May 12, 2026Duc Anh Le, Tien-Phat Nguyen, Thien Huu Nguyen +2Reinforcement LearningRL for Language Model Reasoning