Verified solutions are not equally useful for preparing reasoning models for reinforcement learning (RL). We present a comprehensive study of route diversity, the variation in the sequences of reasoning steps in supervised fine-tuning (SFT) data, and propose a lightweight, rule-based fingerprint to select for it. From one pool at one budget, with matched training recipes and checkpoints, selecting diverse rather than similar routes improves post-RL problem coverage across puzzles and mathematics, including on problems harder than those seen in either training stage. In synthetic experiments, route-diverse SFT improves OLMo3-7B's pass@8 by 16.9 points on environments held out from SFT. In a single-model condition, where one model writes every candidate, diverse selection gains up to 6.2 points of mean pass@8 across 10 mathematics benchmarks. Pre-RL diagnostics suggest why: diverse SFT can produce both successful and failed attempts on more prompts despite slightly lower mean accuracy, giving group-relative RL more prompts with a learning signal. On 3 open-source corpora, our CPU-only selector, without model calls, outperforms more expensive alternatives in every comparison of mean post-RL performance. These results identify reasoning-route diversity as a practical criterion for selecting SFT data that better prepares models for RL.
Figures & tables
Figure 1 : OLMo3-7B on RLVE after the same RL. Diverse (purple circles) solves more problems than Similar (orange open squares) on environments seen and unseen in SFT and in every difficulty band. It solves almost all of what Similar solves, plus 1,133 questions that Similar misses, while Similar solves 53 that Diverse misses (c). (a,b) Gold marks the gap. (c) Solved sets at eight attempts, with gold marking the shared set. Single runs. The generation cap is 16,384 tokens for (b) and the Seen/Unseen 8-sample points in (a), and 32,768 for the All curve, the 32-sample points, and (c).
Figure 2 : At a fixed SFT trajectory budget, more teacher sources give higher post-RL coverage. (a) Held-out Enigmata pass@64 after RL on Enigmata, for Qwen3-1.7B and Qwen3-4B students and SFT pools of 16 and 399 environments, where d is the number of teachers. (b, c) Qwen3-1.7B on OMEGA’s out-of-distribution compositional problems after RL on OMEGA’s training set (b) and reasoning-gym (c).
Figure 3 : At a fixed SFT trajectory budget, twelve teachers beat one after RL on Enigmata (a) and after RL on DAPO-Math-17k (c). (a) Enigmata coverage of Qwen3-1.7B with twelve solutions per prompt from twelve teachers (purple circles) or from one teacher (orange open squares). Gray diamonds mark RL without SFT. (b) Enigmata pass@64 of Qwen3-1.7B after RL for the verified twelve-teacher recipe (purple circles) and the single-Qwen3-14B-teacher recipe (orange open squares), on in-domain (ID) and out-of-domain (OOD) problems. (c) Relative gain of twelve teachers over one ( d=1 ) on AIME 2024 (A’24), AIME 2025 (A’25), MATH-500 (M500) and Minerva (Min.), hatched for the 16-environment pool and solid for the 399-environment pool.
Figure 4 : Topology-based selection. (a) Verified solutions to one prompt x as sequences of step events. (b) Route y3 (gold) as a fingerprint of event frequencies, positions, shape, and transitions. (c) Selection in fingerprint space and at the same budget n=6 , nearest-centroid selection (orange) stays inside the dashed circle and farthest-point selection (purple, numbered in order) spreads out.
Figure 5 : Answer diversity of correct Qwen3 completions on RLVE after the same RL.
Figure 6 : OLMo3-7B reward-signal diagnostics before RL.
Figure 7 : One-teacher route selection. Orange tops: Similar; stack tops: Diverse. Labels: point gain. Axis starts at 28%.
Figure 8 : Relative gain of our selection over each selection baseline in mean score, with both conditions evaluated at the final RL checkpoint, step 64. FPS :farthest-point selection. Topo., Grad., Embed. and Lex. are the topology (our fingerprints with a simpler rule), gradient-diversity, embedding and lexical baselines of Appendix C.4 .
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
x,y
prompt, candidate solution trajectory (route)
V(x,y)
environment verifier, 1 iff y solves x
C+
pool of verified candidate routes
ϕ(y)
topology fingerprint (abstracted transition signature) of route y
d(⋅,⋅)
distance in fingerprint space
Ddiv,Dsim
matched diverse / similar SFT conditions, with ∣Ddiv∣=∣Dsim∣
Table 2 : How each dataset instantiates the shared procedure.
Pool statistic
Diverse pool
Similar pool
Ratio
Response pairwise distance
11
5.0
2.3
Topology pairwise distance
23
9.8
2.3
Topology-vector variance
0.14
0.068
2.0
Pattern entropy
4.9
4.9
1.0
Active-pattern count
32
31
1.0
Appendix
Table 3 : Spread statistics of the 50,000-row diverse and similar RLVE selections (the Qwen3-4B pair at 50,000 rows in Figure 19 ). Per prompt, the two selections use about the same number of reasoning patterns, spread about as evenly, and the distance between solutions is about 2.3 times larger in the diverse selection.
Sweep
Construction
Evaluation
OMEGA source sweep
Qwen3-1.7B; one through five teachers over RLVE pools of 16 and 399 environments; RL on OMEGA’s training set.
RL step 350; 265 OMEGA compositional problems from its out-of-distribution set; 64 samples per problem (Figures 2(b) and 9 ).
Task pool
Qwen3-1.7B, ten reasoning-gym task-pool sizes from 2 to 92 tasks.
Pre-RL Enigmata evaluation, 93 held-out problems and 256 samples per problem. Three seeds per pool size.
Transfer sweeps
Qwen3-1.7B; one through five teachers over RLVE pools of 16 and 399 environments; RL on reasoning-gym tasks.
RL step 350; 265 OMEGA compositional problems and four mathematics benchmarks; 64 samples per problem, temperature 1.0 and an 8,192-token cap (Figures 11 and 13 ).
Qwen3-4B mathematics
One and twelve teachers over RLVE pools of 16 and 399 environments; RL on DAPO-Math-17k.
Four mathematics benchmarks, RL step 200, 64 samples per problem, 8,192-token cap (Appendix B.2 ).
Enigmata teacher grid
Qwen3-1.7B and Qwen3-4B; one, six, or twelve teachers over RLVE pools of 16 and 399 environments; RL on Enigmata.
Held-out Enigmata, RL step 750 (1.7B) and 500 (4B). The 4B evaluation has 486 problems and 64 samples per problem (Figure 2(a) ).
Enigmata decomposition
Qwen3-1.7B; twelve retained solutions per prompt, from one or twelve teachers; RL on Enigmata.
Sampled pass@k at k∈{1,8,32,64} on 125 in-domain problems; direct RL provides a third baseline (Figure 3(a) ).
Appendix
Table 4 : Protocols for the supporting construction sweeps. Teacher count is the number of generators that supply solutions across the corpus.
Figure 9 : Complete OMEGA source sweep behind Figure 2(b) . Marks are post-RL gains in OMEGA compositional coverage, in points, of two- through five-teacher conditions over the one-teacher condition, for SFT pools of 16 and 399 environments. RL trains on OMEGA’s training set. All 56 comparisons are positive. Lines are means and bands observed ranges.
Figure 10 : Held-out Enigmata coverage in percent (shading) for Qwen3-1.7B after SFT on pools of 2 to 92 reasoning-gym tasks. Coverage generally rises, with diminishing gains and local reversals. White rings mark each sampling budget’s best pool. Three-seed means, before RL.
Figure 11 : Full OMEGA compositional profile behind Figure 2(c) , after RL (step 350), with d denoting teacher count. Open marks give every d=2,…,5 coverage difference from d=1 in points at each budget k , for SFT pools 16 (circles) and 399 (squares). RL trains on reasoning-gym tasks, and evaluation uses OMEGA’s out-of-distribution compositional problems, which lie outside both training stages. All 32 comparisons are positive. Lines are means and bands observed ranges.
Figure 12 : Mathematics transfer for Qwen3-1.7B after RL on reasoning-gym tasks, at saved step 350. Five teacher sources are compared with one at equal total SFT trajectory counts, with one training run per condition. Evaluation uses 64 draws per problem, temperature 1.0, and an 8,192-token cap. Bars show relative gain, 100(scored=5−scored=1)/scored=1 in percent. pass@1 is left and pass@64 right. Hatched and solid bars denote pools 16 and 399. A’24, A’25, M500, and Min. denote AIME 2024, AIME 2025, MATH-500, and Minerva. These four benchmarks are not part of either training set. Figure 13 gives every teacher-count level.
Figure 13 : Complete mathematics ladders behind Figure 12 , at (a) pass@1 and (b) pass@64. Each mark is one d=2,…,5 score difference from d=1 in points, for one benchmark and pool (16 and 399). Here d counts teacher sources at a fixed SFT trajectory budget. Scores are measured after RL on reasoning-gym tasks, at RL step 350, and neither SFT nor RL trains on the four benchmarks. All 32 differences are positive at pass@1, and 30 of 32 are positive at pass@64. Marks left of zero favor d=1 .
Figure 14 : Qwen3-1.7B SFT conditions with d=1,…,5 teacher sources at equal trajectory counts over environment pools 16 and 399, then identical GRPO on DAPO-Math-17k mathematics, outside the SFT domain. Each mark is the score difference from d=1 for one benchmark and pool at (a) pass@1 and (b) pass@64.
pass@1 (%)
pass@64 (%)
Benchmark
Pool
1 teacher
12 teachers
1 teacher
12 teachers
AIME 2024
16
16
22
43
50
AIME 2024
399
16
19
47
53
AIME 2025
16
9.5
17
43
53
AIME 2025
399
11
19
40
53
MATH-500
16
34
65
81
83
Appendix
Table 5 : Qwen3-4B mathematics scores in percent, higher is better, at saved checkpoint step 200. Pool is the SFT environment pool. Evaluation uses 64 samples per question, temperature 1.0, and an 8,192-token generation cap. AIME has 30 questions per edition, MATH-500 has 500, and Minerva has 272. There is one training run per condition. pass@1 averages correctness over the 64 draws; pass@64 is the fraction of problems with at least one correct draw.
Figure 15 : Three released corpora, with both conditions at the same RL step. Open squares mark Similar at zero, filled circles Diverse, and bars their signed difference in points. (a) Mean sampled accuracy, with strict prompt accuracy for IFEval and IFBench. (b) pass@8 for the seven eight-response benchmarks. (c) pass@4 for OMEGA and Enigmata. All rows use RL step 8 for Nemotron-Cascade 2 and OpenThoughts3, and step 16 for INTELLECT-3. On INTELLECT-3 Enigmata, Diverse reaches 1.75% pass@4, compared with 6.00% for Similar.
Figure 16 : OpenThoughts3: relative gains over each selection baseline on individual benchmarks. Purple indicates positive gains, orange losses, and white zero. The pass@1 and pass@8 panels show every reported benchmark cell.
Figure 17 : INTELLECT-3: benchmark-level relative gains over each selection baseline, on the same scale as Figure 16 . Ties and losses remain visible alongside improvements.
Figure 18 : Nemotron-Cascade 2: benchmark-level relative gains over each selection baseline, on the same scale as Figure 16 . All three evaluated approaches are shown.
Figure 19 : Qwen3 students on RLVE after the same RL: relative gain of Diverse over Similar SFT. Diverse leads on every metric at both model sizes and both selection budgets. The right column gives the Similar and Diverse scores in percent. pass@1 denotes mean sampled accuracy, and pass@32 and pass@64 denote sampled coverage.
Student
Prefix
Questions
Diverse
Similar
Relative gain (95% CI)
Qwen3-4B
128
861
0.74
0.64
17% ( 14 , 19 )
512
675
0.82
0.77
6.0% ( 5.2 , 6.8 )
Qwen3-1.7B
128
1,327
0.57
0.49
15% ( 12 , 18 )
512
1,258
0.74
0.71
4.0% ( 3.3 , 4.7 )
Appendix
Table 6 : Mean pairwise bigram Jaccard distance among correct completions. Prefix lengths are whitespace-token counts. Confidence limits are percentages for the relative gain, from paired resampling of evaluation environments.
Figure 20 : (a) Mean sampled accuracy and (b) coverage in percent by RLVE difficulty, after identical RL from diverse (purple circles) and similar (orange squares) SFT. Difficulties 1 to 10 use 32 samples. The shaded range, 11 to 15, uses 64. Qwen3-4B-Base with 200,000 SFT rows selected from the shared pool. Mean sampled accuracy uses the scored question set, and coverage uses its programmatic-reference-answer subset. This pair is separate from the Qwen3-4B 200,000-row pair in Figure 19 .
Figure 21 : Post-RL mathematical and executable-route comparisons. Diverse leads Similar at pass@64 in every panel, and in (b,c) the gap grows with k . (a) Relative OMEGA pass@64 gains, 100(condition−Similar)/Similar . Purple bars show Diverse, the orange zero line denotes Similar, and gray diamonds show direct RL without SFT on the same relative scale. For Qwen3, Diverse and Similar denote multi-model and single-model corpora; for OLMo3-7B, they denote topology-selected subsets of one shared pool. (b,c) OLMo3-7B on program simulation and Sokoban at k=1,4,64 . In both, k=1 is sampled pass@1. Both panels estimate pass@4 from 64 samples per problem and report empirical coverage at k=64 (Appendix A.3 ). Scores are percentages. In (b), Diverse is the multi-model corpus and Similar the single-model corpus of equal size. In (c), both are selected from one pool. Purple circles mark Diverse and orange open squares Similar. Both conditions in each comparison use RL step 75. One run per condition.
RLVE
Candidates and verification
Procedurally generated reasoning-gym environments, each with a rule-based verifier. Each trace gets a 1,373-dimensional lexical-topological fingerprint with 159 continuous, 1,150 topology and 64 pattern features.
Selection conditions
Similar sets use nearest-centroid selection. Diverse sets cluster the fingerprints and use greedy farthest-point selection inside each cluster, with budgets in proportion to cluster size. At both Qwen3 sizes, both conditions select from the same shared pool at 50,000 and 200,000 rows. A second Qwen3-4B pair at 200,000 rows appears in Figure 20 . Both OLMo3-7B conditions retain 50,000 rows from a shared 64-environment pool.
Evaluation
One fixed evaluation set spanning 384 environments and difficulties 1 to 15. Qwen3 runs report pass@32 at difficulties 1 to 10 and pass@64 at 11 to 15, over the 3,287 and 1,587 questions with a programmatic reference answer. The OLMo3-7B run reports pass@8 and pass@32 over 5,682 questions, 937 from the 63 SFT environments in the set and 4,745 from the 321 environments held out from SFT. The shared RL pool spans all 384 environments. Metrics and diagnostic caps are given in Appendix A.3 .
Training recipe
Full-parameter SFT for one epoch, learning rate 10−5 , 8,192-token sequences. GRPO uses difficulties 1 to 10 for both Qwen3 and OLMo3-7B: 75 steps, 8 rollouts, actor learning rate 10−6 , KL coefficient 0.001 , prompt 4,096 and response 16,384 tokens, at 128 prompts per step.
OMEGA
Candidates and verification
Candidates are kept when the OMEGA answer checker accepts them. Single-model traces come from Qwen3-4B. Multi-model traces continue each response across a roster of open reasoning models and are kept in English (Appendix A ). Topology selection uses a 323-dimensional strategy-step fingerprint over an accepted pool of 575,699 traces, 416,727 multi-model and 158,972 single-model.
Appendix
Table 7 : Setup for the RLVE, OMEGA, Sokoban, program-simulation, real-data and single-teacher testbeds. Each block gives candidate generation and verification, the diverse and similar conditions, the evaluation, and the SFT-to-GRPO training recipe.
Recent advances in large language models (LLMs) have demonstrated that reinforcement fine-tuning of pretrained base models can lead to significant gains in reasoning performance at inference time. In this work, we theoretically analyze why reinforcement fine-tuning induces better reasoning ability than purely supervised fine-tuning (SFT) methods. We model chain-of-thought (CoT) reasoning as a pathfinding problem on graphs and compare the popular method of reinforcement learning with verifiable rewards (RLVR) against traditional SFT. We prove that SFT, when trained on golden shortest paths without negative examples, fails to learn how to efficiently backtrack. In contrast, an RLVR-trained model can learn how to efficiently backtrack from dead ends using only outcome reward. This leads to an exponential separation in inference-time compute between the two methods, and demonstrates that RLVR leads the model to learn the location of difficult decisions in a reasoning chain, ultimately allowing for better allocation of inference-time compute. Finally, we show that the reasoning traces of an RLVR model can be distilled to train a base model to backtrack efficiently as well.
Stanley Wei, Juno Kim
1Princeton University · University of California, Berkeley.
Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterparts on mathematical reasoning tasks; Yet the mechanistic basis for this advantage remains unclear. We therefore ask, what internal representational differences enable RL models' superior performance? Our work presents two converging lines of evidence: First, linear probes trained on layer-wise hidden states reveal that RL models tend to achieve higher accuracy in predicting answer correctness compared to SFT models, indicating more linearly separable and structured representations. Second, mean ablation studies show that RL models develop a hierarchical architecture where deeper layers become progressively more critical, whereas SFT models distribute importance uniformly across layers. Together, these findings demonstrate that RL training fundamentally restructures how models represent and process reasoning problems. Finally, we analyze token-count variability under repeated sampling across problems to assess adaptive compute allocation. While we observe higher variability in some RL-tuned models than in their SFT counterparts, we see strong consistency in others, suggesting that token allocation may depend more on the overall training pipeline than on RL versus SFT alone. We believe this token-allocation variability reveals the spread of plausible on-policy reasoning, highlighting which models exhibit stable policies versus those that are under-determined, potentially non-identifiable solution behaviour.
The standard post-training recipe for large reasoning models, supervised fine-tuning followed by reinforcement learning (SFT-then-RL), may limit the benefits of the RL stage: while SFT imitates expert demonstrations, it often causes overconfidence and reduces generation diversity, leaving RL with a narrowed solution space to explore. Adding entropy regularization during SFT is not a cure-all; it tends to flatten token distributions toward uniformity, increasing entropy without improving meaningful exploration capability. In this paper, we propose CurioSFT, an entropy-preserving SFT method designed to enhance exploration capabilities through intrinsic curiosity. It consists of (a) Self-Exploratory Distillation, which distills the model toward a self-generated, temperature-scaled teacher to encourage exploration within its capability; and (b) Entropy-Guided Temperature Selection, which adaptively adjusts distillation strength to mitigate knowledge forgetting by amplifying exploration at reasoning tokens while stabilizing factual tokens. Extensive experiments on mathematical reasoning tasks demonstrate that, in SFT stage, CurioSFT outperforms the vanilla SFT by 2.5 points on in-distribution tasks and 2.9 points on out-of-distribution tasks. We also verify that exploration capabilities preserved during SFT successfully translate into concrete gains in RL stage, yielding an average improvement of 5.0 points.
Hao Wang, Hao Gu, Hongming Piao +6
1City University of Hong Kong · 2The Hong Kong University of Science and Technology · 3The Chinese University of Hong Kong