A verified reference solution provides a correct trajectory for training a reasoning model. Alternatively, a prefix of the reference can guide the model in generating a trajectory of its own. How much reference guidance should we provide? We study this question through prefix continuation, where the model continues from a reference prefix and keeps the resulting trajectory if it passes verification, falling back to the reference otherwise. Since both procedures produce correct trajectories, we compare their distributions with the ideal distribution, the model's own distribution conditioned on successful verification. For one continuation, we derive the KL divergence in closed form, which, up to a bounded term, decreases with the product of the probability of generating a different correct trajectory and the reference surprisal, the negative log probability of the reference suffix given the prefix. Since a longer prefix tends to raise the former but lowers the latter, continuation success alone does not determine the preferred amount of guidance. From this analysis, we learn a prefix selector shared across training questions from continuation outcomes, without estimating success probabilities or additional generation. The resulting Adaptive Reference Guidance (ARG) constructs correct trajectories within a fixed generation budget, and we apply it to all-failure groups in Group Relative Policy Optimization (GRPO). Experiments on Qwen3-4B and Qwen3-8B across five mathematical reasoning benchmarks show that ARG achieves the highest aggregate pass@12 among the evaluated methods with competitive average sampled accuracy.
Figures & tables
Model
Method
AIME24
AIME25
AIME26
HMMT_Feb
HMMT_Nov
Avg.
Qwen3-4B
Base
50.00
43.33
43.33
23.33
33.33
38.67
SFT
63.33
63.33
66.67
36.67
36.67
53.33
GRPO
76.67
70.00
73.33
53.33
46.67
64.00
OPSD
56.67
36.67
40.00
40.00
40.00
42.67
SAGE
73.33
70.00
70.00
43.33
56.67
62.67
BREAD
80.00
73.33
73.33
50.00
56.67
66.67
Table 1: Main experimental results in pass@12 percentages. Avg. is the mean over the five benchmarks. Within each model, bold and underline mark the best and second-best values in each column.
Table 2: Ablation results averaged over the five benchmarks. Resampling uses eight continuations from the empty prefix without reference fallback. The weighting ablation removes St/S0 from the selector feedback.
Figure 1: Training dynamics of a Qwen3-8B ARG run. Panel (a) reports successful replacement rates, panel (b) reports mean selector feedback Wt , and panel (c) shows the selection probabilities. In panels (a) and (b), a point at step s summarizes all continuations at that level across the past 50 steps, including failures and candidate trajectories not accepted for training. Panel (c) shows the state after each training step without smoothing.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
ARG configuration
Questions per update
32
Ordinary responses per question G
8
Continuations per all-failure group K
8
Guidance levels A
{1,1/2,1/4}
Total acceptance cap in Algorithm 1
4
Maximum accepted trajectories with empty prefixes
4
Appendix
Table 3: ARG configuration for both model sizes.
Setting
GRPO-based runs
SFT
OPSD
Learning rate
10−6
10−6
5×10−7
Warmup updates
10
10
0
Weight decay
0.01
0.01
0
Gradient norm clipping
1
1
0.1
Responses per question
8
Reference
1
Response length limit
8192
8192
1024
Appendix
Table 4: Recorded configurations for both model sizes. GRPO-based runs here comprise GRPO, SAGE, BREAD, LUFFY, ARG, resampling, and the weighting ablation. The reference-insertion ablation is specified by its replacement rule in Section 5.3 .
Method
4B checkpoint
8B checkpoint
Base
0
0
SFT
25
25
OPSD
60
100
GRPO, SAGE, BREAD, LUFFY, ARG
200
200
Resampling
200
N/A
w/o weighting
200
200
Appendix
Table 5: Reported checkpoint steps for both model sizes. Base uses the released parameters. SFT and OPSD use benchmark-selected checkpoints, and the remaining trained entries report step 200.
Setting
Evaluation value
Temperature
1
Top- p
0.95
Top- k
20
Maximum new tokens
16,384
Maximum context length
32,768
Scored responses per question
12
Appendix
Table 6: Evaluation configuration shared by all reported entries.
Model
Method
AIME24
AIME25
AIME26
HMMT_Feb
HMMT_Nov
Avg.
Qwen3-4B
Base
22.78
22.22
13.33
11.94
9.17
15.89
SFT
19.44
20.56
16.94
9.72
9.17
15.17
GRPO
47.22
42.22
46.67
26.94
29.17
38.44
OPSD
22.22
18.89
20.83
14.44
13.06
17.89
SAGE
44.17
39.17
40.83
25.00
28.06
35.44
BREAD
56.39
47.78
47.78
29.17
35.56
43.33
Appendix
Table 7: Main experimental results in avg@12 percentages. Avg. is the mean over the five benchmarks. Within each model, bold and underline mark the best and second-best values in each column.
Model
Method
AIME24
AIME25
AIME26
HMMT_Feb
HMMT_Nov
Avg.
Qwen3-4B
Reference insertion
70.00
60.00
63.33
43.33
46.67
56.67
Resampling
76.67
63.33
63.33
43.33
53.33
60.00
w/o weighting
76.67
70.00
66.67
50.00
53.33
63.33
ARG
80.00
73.33
73.33
43.33
66.67
67.33
Qwen3-8B
w/o weighting
83.33
70.00
66.67
50.00
50.00
64.00
ARG
83.33
73.33
73.33
56.67
63.33
70.00
Appendix
Table 8: Benchmark-level ablation results in pass@12 percentages. Avg. is the mean over benchmarks. Here w/o weighting denotes ARG without surprisal weighting.
Model
Method
AIME24
AIME25
AIME26
HMMT_Feb
HMMT_Nov
Avg.
Qwen3-4B
Reference insertion
36.67
34.17
34.17
17.78
19.44
28.44
Resampling
44.44
35.56
40.00
21.67
28.33
34.00
w/o weighting
43.61
40.83
39.17
21.39
28.33
34.67
ARG
52.50
46.67
46.39
26.94
33.06
41.11
Qwen3-8B
w/o weighting
45.28
33.89
38.33
18.33
23.61
31.89
ARG
55.28
46.94
46.39
27.22
35.56
42.28
Appendix
Table 9: Benchmark-level ablation results in avg@12 percentages, computed from the same scored responses as Table 8 . Avg. is the mean over benchmarks.
Figure 2: Proportion of all-failure groups that accept candidate trajectories. A point at step s aggregates steps s−49 through s . Reference fallback is excluded. The curves come from their respective training runs.
Figure 3: Selection probabilities of the prefix selector with and without surprisal weighting. Curves show the state after each training step without smoothing. Steps with no selector update retain the previous probabilities. The level α=1 uses an empty prefix, while smaller levels leave less of the reference surprisal to the continuation.
Figure 4: Training dynamics of a Qwen3-4B ARG run. In panels (a) and (b), a point at step s averages the continuations at each level from steps s−49 through s . Panel (c) shows the selection probabilities after each training step without smoothing. The colors and corresponding axis ranges match Figure 1 .
Figure 5: Sources of accepted trajectories among all-failure groups in consecutive windows of 50 training steps. The mixed category contains groups that accept both guided and unguided candidate trajectories. Fallback supplies the complete reference when no candidate trajectory is accepted. Counts below the windows give the number of all-failure groups.
Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot generate any correct solutions, it receives \textit{zero} learning signal. Providing privileged guidance during training, such as solution prefixes, can help overcome this learning cliff by steering the model towards {correct solutions with non-zero reward}. {We call these rollouts \textit{off-context}: they are generated from a training prompt that contains privileged guidance, while the target objective is defined by the original prompt without that guidance.} {We introduce} Off-Context GRPO (OC-GRPO), a minimally modified variant of GRPO that uses guided rollouts but applies an importance-corrected objective to steer the update back toward the original unguided objective, avoiding the mismatch that destabilizes uncorrected guided training. Empirically, our algorithm achieves a 3.9% absolute improvement (13.8% relative gain) over vanilla GRPO on average across standard mathematical reasoning benchmarks with negligible additional cost.
Reinforcement Learning with Verifiable Rewards (RLVR) for language-model reasoning can fail at both extremes of task difficulty: easy prompts often produce all-correct, low-diversity rollout groups with little gradient signal, while hard prompts can produce all-incorrect groups with no positive reward. We introduce ExTra (Exploratory Trajectory Optimization), a GRPO-compatible framework that extracts exploration signals from the model's own rollouts. ExTra combines two mechanisms: (i) a novelty reward that adds embedding-based diversity bonuses after GRPO normalization, rewarding diverse correct solutions; and (ii) entropy-guided prefix regeneration, which scores partial trajectories using entropy signals and continues exploration from promising intermediate steps. Across six mathematical reasoning benchmarks, ExTra improves Qwen3-1.7B over GRPO by about +5 points on pass@1 and +7 points on pass@16, showing that trajectory-level exploration signals can improve both single-sample accuracy and inference-time coverage.
Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT adds a repair update for those failures without changing rollout generation: it derives a plan from the reference solution, compares token probabilities with and without that plan, and gives higher weight to tokens that become more predictable under plan conditioning. This turns all-wrong prompts into selective, structure-aware learning signals instead of uniform imitation. Across three backbones and eight reasoning benchmarks, SORT improves over GRPO and guidance baselines, with largest gains on weaker models.
Duc Anh Le, Tien-Phat Nguyen, Thien Huu Nguyen +2
Independent Author · Hanoi University of Science and Technology, Hanoi, Vietnam · University of Oregon, Eugene, OR, USA +1