Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints. A key promise of RLVR is the discovery of new reasoning strategies. In principle, a model can sample novel ideas absent from its prior training data. In practice, however, augmenting RLVR with strong novelty incentives has seen limited success and can degrade model quality. Because verifiable rewards supervise only a narrow slice of the model's knowledge and behavior, such degradations are difficult to recover from. Instead, we decouple exploration from optimization in a framework we call Exploration-Distillation (ExpDis). We train one or more explorer policies with a novelty bonus in the reward, filter their trajectories for correctness and quality, and distill them into a separate student policy. The student policy is then trained without a novelty bonus. We repeat the above procedure for several rounds, alternating between exploration and optimization. This decoupling allows us to aggressively scale exploration without degrading the student policy. Across seven mathematical reasoning benchmarks and two model families, ExpDis outperforms DAPO at the same wall-clock budget. Moreover, we observe improved pass@k scaling, indicating that ExpDis produces models that generate more diverse correct solutions.
Figures & tables
Figure 1 : Decoupling exploration from optimization improves math accuracy without degrading prior capabilities. Left: Rather than coupling exploration and optimization in a single RLVR run, ExpDis assigns these roles to separate policies. We first train an explorer μ with a novelty bonus in its reward, then distill a diverse pool of its reasoning traces to cold start a student, and finally train the student with standard RLVR using only a correctness reward. As depicted in the figure, this allows us to explore a diverse set of reasoning strategies and preserve this diversity in our student policy. (a) ExpDis outperforms DAPO in both average accuracy and pass@ k , even when DAPO is trained for 4× as long (Qwen3-4B, mean pass@ k over five math benchmarks). (b) Adding a novelty bonus (RND) directly to the DAPO reward degrades performance on prior capabilities relative to the base model. Because the reward supervises only a narrow slice of the model’s knowledge (here, mathematical reasoning), achieving high reward does not prevent general capability degradation. By decoupling the explorer from the student policy, we avoid this degradation.
Figure 2 : ExpDis scales across parallel explorers and iterations. Within a round, several explorers are trained in parallel with a novelty bonus, and their traces are pooled to distill a single student. Across rounds, that student initializes both the next round’s explorers and the next round’s student, so every round starts from a stronger model. Each round trains on its own shard of the training prompts, and the novelty weight anneals from λ1=0.75 to λ4=0.25 . We study how to allocate a fixed RL budget across explorers and rounds in Section 5 .
Figure 3 : ExpDis outperforms all baselines under the same RL compute budget. By partitioning this budget into alternating rounds of exploration and optimization (as described in Section 3.2 ), we can surpass DAPO and even a single round of ExpDis. We report full benchmark results in Appendix B .
Figure 4 : Under a fixed RL compute budget, how we allocate compute impacts downstream student policy performance. We vary the number of explorers and rounds under a fixed compute budget, and find complementary impacts of each scaling axis. Under a fixed budget, adding rounds improves pass@1 more than adding explorers, and combining both gives the largest gains in pass@1 and pass@64. All results are reported on Qwen3-1.7B.
Figure 5 : Explorer policy can be driven to more diversity without harming the student. Left: We study how explorer policy diversity impacts the downstream student in ExpDis. In general, we find that we can push the diversity of the explorer policy without corrupting the downstream student policy. Since we filter for correctness and quality before distilling, a moderate increase in explorer diversity (for example, lexical diversity) brought on by tuning the novelty bonus weight λ does not negatively impact the student policy. Similarly, we can increase the explorer token entropy far beyond that of DAPO (+novelty) without incurring downstream reasoning degradation. Right: We find that ExpDis generations contain more semantically diverse mathematical reasoning strategies. We define each diversity metric in Appendix D .
Figure 7 : In successive ExpDis iterations, the student captures more of the explorer’s diversity while avoiding its degradations. We track the explorers and the student as they evolve during several ExpDis rounds and report the final baseline policies to the right of each panel. (a) Explorer RL causes a rise in token entropy often at the expense of accuracy, whereas the student’s entropy generally rises together with its accuracy across rounds. (b) The explorer policy semantic diversity stays roughly constant, and the student closes most of the gap to it by the final round. (c) We find that the student does not retain the lexical diversity of its explorers: it declines over rounds even as the student’s accuracy improves. Together with Figure 5 , where large λ harms the student, this suggests that lexical diversity is not a useful target for exploration. (d) Reasoning faithfulness, as defined by Rahman et al. [31] , measures whether a response’s reasoning supports its final answer. We find that novelty bonus typically causes faithfulness to decline, though the impact is mild by the final round. Results are averaged over three models (breakdown in Figure 14 ).
explorer 200, student 100, baselines 300, DAPO ( 4× steps) 1,200 total
Max completion length
32,768
Soft overlong penalty
linear over the final 6,554 tokens
Optimizer
AdamW ( β=(0.9,0.95) , no weight decay)
Appendix
Table 1 : Hyperparameters. Shared by all models and all runs unless noted.
Figure 8 : Training dynamics of continued training on Qwen3-1.7B with our ExpDis setup ( λ=0.5 ).
Configuration
R×K
Explorer updates
Student updates per round
Single-Explorer
1×1
200
100
Breadth K=2
1×2
100 each
100
Breadth K=3
1×3
67/67/66
100
Breadth K=5
1×5
40 each
100
Breadth K=7
1×7
29/29/29/29/28/28/28
100
Depth R=4
4×1
50 per round
25
Appendix
Table 2 : Update allocation per configuration. MR-ME (multi-round, multi-explorer) combines both axes: R rounds with K explorers per round. Every ExpDis run uses 200 explorer and 100 student updates (19,200 selected rollouts); single-model baselines spend all 300 on one model. Parallel explorers run concurrently, so wall-clock time is equal across rows. DAPO ( 4× steps) is the only exception.
Training rollouts
Evaluation
Temperature
1.0
0.6
top- p / top- k / min- p
0.95 / 20 / –
0.95 / 20 / 0
Max completion tokens
32,768
32,768
Samples per problem
16
64 ∗
Appendix
Table 3 : Sampling parameters. ∗ 32 for AMC23 and 8 for GSM8K.
Figure 9 : Full pass@ k results.
Model
Method
@1
@2
@4
@8
@16
@32
@64
Qwen3-1.7B
Base
43.42
52.18
58.26
63.02
66.95
70.63
73.82
Qwen3-1.7B
DAPO
45.63
54.31
60.67
65.54
69.16
72.14
74.64
Qwen3-1.7B
DAPO ( 4× steps)
47.28
56.44
62.83
67.41
70.80
73.62
76.24
Qwen3-1.7B
DAPO (+novelty)
45.08
53.78
60.07
64.91
68.61
71.77
74.43
Qwen3-1.7B
ExpDis (single-round)
48.93
58.55
64.98
69.27
72.43
75.10
77.83
Qwen3-1.7B
ExpDis
51.81
61.40
68.03
72.94
76.97
80.29
83.04
Appendix
Table 4 : Mean pass@ k over the five primary benchmarks.
Figure 11 : On Qwen3-1.7B, we find that a novelty bonus weight of λ =0.5 on the explorer policy roughly maximizes downstream student policy performance, though several values of λ outperform the DAPO baseline.
Figure 12 : We report the degradation from adding a novelty bonus directly to the DAPO reward.
Method
rnovelty
AIME24
AIME25
AIME26
MATH500
Minerva
Mean
DAPO
–
50.05
36.93
37.66
74.24
29.27
45.63
DAPO ( 4× steps)
–
51.41
38.05
41.96
75.02
29.98
47.28
ExpDis (ours)
RND
52.76
39.17
46.25
75.80
30.69
48.93
ExpDis
kNN
50.83
37.50
43.75
75.00
29.75
47.37
ExpDis
Elliptical
51.56
38.23
44.79
75.30
30.40
48.06
Appendix
Table 5 : We ablate the choice of novelty metric used as our novelty bonus.
Figure 13 : Impact of increasing rounds given a fixed number of explorers per round.
Model
Entropy
Semantic
Ans. entropy
Distinct ans.
InterDistinct-4
Base
0.237
0.100
1.42
8.94
0.311
DAPO
0.251
0.096
1.37
8.50
0.308
DAPO ( 4× steps)
0.265
0.091
1.35
8.46
0.302
ExpDis (single-explorer)
0.306
0.120
1.41
9.93
0.357
ExpDis
0.391
0.143
1.30
9.44
0.339
Appendix
Table 6 : We report diversity metrics for Qwen3-1.7B models on AIME24 generations.
Explorer
Student
Model
Method
Inter4
Ans@ n
Inter4
Ans@ n
Qwen3-1.7B
Base
–
–
0.311
8.94
Qwen3-1.7B
DAPO
–
–
0.308
8.50
Qwen3-1.7B
DAPO (+novelty, λ=0.5 )
–
–
0.342
9.39
Qwen3-1.7B
ExpDis ( λ=0.25 )
0.403
11.60
0.346
9.63
Qwen3-1.7B
ExpDis ( λ=0.5 )
0.422
12.61
0.357
9.93
Appendix
Table 7 : We report InterDistinct-4 (Inter4) and AnswerDistinct@ n (Ans@ n ) for the explorer and the student on AIME24 across explorer weights λ .
Figure 14 : We study the evolution of student and teacher policies during several stages of ExpDis.
Figure 21
Method
Dataset
AIME24
AIME25
AIME26
MATH500
AMC23
Minerva
GSM8K
GRPO
DAPO-Math-17K
47.81
35.36
35.68
72.84
83.75
28.02
89.89
GRPO
DeepScaleR-17K
47.96
35.52
35.68
72.71
84.09
27.86
89.81
Dr. GRPO
DAPO-Math-17K
48.49
35.99
36.46
73.34
84.38
28.45
90.04
Dr. GRPO
DeepScaleR-17K
48.83
36.26
36.51
73.18
84.78
28.20
89.95
DAPO
DAPO-Math-17K
50.05
36.93
37.66
74.24
85.16
29.27
90.25
DAPO
DeepScaleR-17K
49.68
37.21
37.76
74.08
85.62
28.90
90.14
Appendix
Table 9 : We ablate the choice of training dataset. We see roughly the same performance across all datasets.
Figure 16 : We ablate the choice of training dataset, sampling a fixed prompt budget from each and training with a single-explorer ExpDis approach. We find a consistent gain over several RLVR baselines.
Reinforcement learning with verifiable rewards (RLVR) has emerged as a scalable paradigm for improving the reasoning capabilities of large language models. However, its effectiveness is fundamentally limited by exploration: the policy can only improve on trajectories it has already sampled. While increasing the number of rollouts alleviates this issue, such brute-force scaling is computationally expensive, and existing approaches that modify the optimization objective provide limited control over what is explored. In this work, we propose NudgeRL, a framework for structured and diversity-driven exploration in RLVR. Our approach introduces Strategy Nudging, which conditions each rollout on lightweight, strategy-level contexts to induce diverse reasoning trajectories without relying on expensive oracle supervision. To effectively learn from such structured exploration, we further propose a unified objective, which decomposes the reward signal into inter- and intra-context components and incorporates a distillation objective to transfer discovered behaviors back to the base policy. Empirically, NudgeRL outperforms standard GRPO with up to 8 times larger rollout budgets, while outperforming oracle-guided RL baseline on average across five challenging math benchmarks. These results demonstrate that structured, context-driven exploration can serve as an efficient and scalable alternative to both brute-force rollout scaling and feasibility-oriented methods based on privileged information. Our code is available at https://github.com/tally0818/NudgeRL.
Reinforcement Learning with Verifiable Rewards (RLVR) for language-model reasoning can fail at both extremes of task difficulty: easy prompts often produce all-correct, low-diversity rollout groups with little gradient signal, while hard prompts can produce all-incorrect groups with no positive reward. We introduce ExTra (Exploratory Trajectory Optimization), a GRPO-compatible framework that extracts exploration signals from the model's own rollouts. ExTra combines two mechanisms: (i) a novelty reward that adds embedding-based diversity bonuses after GRPO normalization, rewarding diverse correct solutions; and (ii) entropy-guided prefix regeneration, which scores partial trajectories using entropy signals and continues exploration from promising intermediate steps. Across six mathematical reasoning benchmarks, ExTra improves Qwen3-1.7B over GRPO by about +5 points on pass@1 and +7 points on pass@16, showing that trajectory-level exploration signals can improve both single-sample accuracy and inference-time coverage.
Reinforcement Learning with Verifiable Rewards (RLVR) has become a widely adopted technique for enhancing the reasoning ability of Large Language Models (LLMs). However, the effectiveness of RLVR strongly depends on the capability of base models. This issue arises because it requires the model to have sufficient capability to perform high-quality exploration, which involves both effectiveness and diversity. Unfortunately, existing methods address this issue by imitating expert trajectories, which improve effectiveness but neglect diversity. To address this, we argue that the expert only needs to provide guidance only at critical decision points rather than the entire reasoning path. Based on this insight, we propose MENTOR: Mixed-policy Expert Navigation for Token-level Optimization of Reasoning, a framework that provides expert guidance only at critical decision points to perform effective and diverse exploration in RLVR. Extensive experiments show that MENTOR enables models capture the essence of expert strategies rather than surface imitation, thereby performing high-quality exploration and achieving superior overall performance. Our code is available online.
Zishang Jiang, Jinyi Han, Tingyun Li +7
School of Data Science, Fudan University · Shanghai Institute of Artificial Intelligence for Education, East China Normal University · College of Computer Science and Artificial Intelligence, Fudan University +1