Probe with Participation Trophies: Random-Reward RL as a Probe of LLM Capability
Authors: Yu Mao, Lei Yu, Zining Zhu, Yusheng Zheng, Haohang Li, Freda Shi, Yutong Yin, Zhaoran Wang, +1 more
Organizations: University of Toronto · Stevens Institute of Technology · University of California, Santa Cruz · University of Waterloo · Vector Institute · Northwestern University
We connect the spurious-reward paradox to a model's reachability and propose random-reward reinforcement learning (RL) as a useful tool for the probing enterprise, addressing a decade-long debate over what probing performance actually reveals about a model. There are two prevailing explanations for the surprising finding that even random rewards can improve the performance of large language models (LLMs): one attributes the gains to particular mechanisms within RL training; the other to data contamination. Our results motivate a different view: spurious-reward RL can probe a model's reachability, or what further training can attain from its current state under specified constraints, beyond what is reflected in its current performance. Two OLMo checkpoints with the same accuracy on synthetic arithmetic (3.5%), for example, reach 8.5% and 55% in their best runs under the same correctness-rewarded RL. Examining OLMo checkpoints across pre-training and mid-training reveals three distinct regimes of training response: early on, RL produces little improvement even when correct answers are rewarded; later in pre-training, rewarding correct answers becomes effective while random rewards remain weak; and, upon entering mid-training, even random rewards can produce large gains. A similar ordering appears in a number-masked supervised fine-tuning (SFT) analysis of these checkpoints, suggesting that the pattern is not specific to a particular RL mechanism. Moreover, RL with random rewards offers a distinctive perspective on what training can make an LLM do, since its reward signal supplies no information about which answers are correct. By asking what training can attain without correctness feedback, it addresses the label-leakage side of a central problem in decodability-based probing: whether a successful probe reveals the model's capabilities or learns the task itself.
Figures & tables
Figure 1 : Development of LLM reachability across pre-training and mid-training (OLMo-2, synthetic GSM). Early checkpoints barely improve under either reward (dormant), later ones improve under GT rewards (receptive), and in mid-training some random-reward runs also reach high accuracy (autodidactic). Each line is a seed ending at its held-out accuracy after 500 RL steps, higher closer to the star; bars give the best across seeds, and (ΔGT,Δrand) marks a gain (+) or none (0). The landscape is schematic; overlapping endpoints (panel I; GT in panel III) are spread slightly.
Response
Learner
Initial
Terminal
Mean
source
score
max
gain
P1
P1
2.5
9.5
+0.6
P2
13.5
80.0
+4.5
P2
P1
2.5
14.0
+2.9
P2
13.5
57.0
+13.7
Table 1: Identical rollouts and random rewards yield different outcomes across OLMo-2 checkpoints. P1: end of pre-training; P2: +5B mid-training. Each row summarizes 16 runs, with terminal results at step 500. Scores are accuracies (%); gains are changes from initial accuracy (percentage points).
Figure 2 : With the same rollouts and rewards, P2 gains more on average and varies more. Panels split runs by rollout source, the x-axis by learner. Dots are the 16 runs per group and lines their means; changes are step-500 minus initial accuracy (percentage points).
Figure 3 : Inputs used to separate task cues from correctness information.
Figure 4 : Random-reward training of OLMo-2 P2 +50B on sequences of randomly sampled tokens: (a) accompanied by the math instruction and (b) presented alone. Panel (c) shows terminal accuracy after 500 training steps. Thin lines and dots show individual runs (32 seeds per condition); highlighted lines show the highest accuracy across seeds at each step. The dashed line marks the initial accuracy of 20.5%. Further comparisons appear in Appendix 11.5 .
Figure 5 : Three representative OLMo-2 checkpoints illustrate the dormant, receptive, and autodidactic stages of training response on synthetic GSM. Top: GT, four seeds per checkpoint; bottom: Random, eight. Thin lines show sampled rollout accuracy (trailing 20-step mean); black lines show pointwise maxima across seeds; dashed lines show frozen greedy held-out accuracy.
Figure 6 : Both RL and SFT show the same three-stage developmental pattern across OLMo-2 checkpoints (synthetic GSM). (a) Frozen and maximum terminal accuracy after 500 RL steps; faint points show individual seeds. (b) Initial and best reported SFT accuracy. Stage bands in both panels follow the RL results in (a). See Appendix 11.4 for the SFT comparison.
Figure 7 : OLMo-2 results on code (a) and lookup (b). The overall trend is similar, although the timing of the transitions is idiosyncratic to each task. Lines show frozen accuracy and maximum accuracy across seeds after 500 RL steps; faint points show individual runs. On code, GT and Random gains first appear together in P2. The dashed line marks chance accuracy on lookup (25%).
Synthetic GSM
Code
Invented-fact lookup
P1 / P2 checkpoints
25 / 12
8 / 4
8 / 4
Training / evaluation problems
800 / 200
764 / 200
800 / 200
GT seeds, P1 / P2
4 / 4
1 / 4
2 / 2
Random seeds, P1 / P2
8 / 8
2 / 8
8 / 8
Response cap (tokens)
96
256
96
Training steps
500
500
500
Table 2: Checkpoint-scan settings. Seed counts are per checkpoint. GT and Random use reference coefficients 0 and 0.01, respectively, in addition to different rewards. Appendix 10.8.5 defines the reference term.
Arm
Training content
Math template
Seeds
B
Original GSM question
Yes
32 (8 reused)
C
Word-shuffled GSM question
Yes
16
A
Unrelated MMLU question stem
Yes
16
D
Random vocabulary tokens
Yes
32
E
Random-token complete prompt
No
32
Table 3: The five training-input conditions. All use random rewards and are evaluated on the same GSM questions. The math template contains the instruction and the Question: and Answer: markers.
ID
Training position
Training step
SmolLM2-1.7B
S1
0.26T
step-125000
S2
5.77T
step-2750000
S3
9.96T
step-4750000
S4
10.75T
step-5125000
OLMo-3-7B
Table 4: SmolLM2 and OLMo-3 checkpoints used in the developmental comparison. P2 token counts are additional to P1.
configuration
reward
reference coefficient
G
r_g4 / r_g8
independent random
0.01
4 / 8
gt_g4
ground truth
0
4
gt_K50
ground truth for 50 steps, then random
0.01
8
anti_K50
inverted ground truth for 50 steps, then random
0.01
8
Table 5: Group-size and steering conditions in the earlier 55-run study. All use learning rate 3×10−5 , 500 updates, and evaluation every 100 steps. Steering supplies correctness-dependent rewards for the first 50 steps.
Batch
Runs
Flagged steps / all steps
Runs with a flagged step
Completed checkpoint scan
636
519 / 318,000
155
Expanded input controls
152
492 / 76,000
112
SmolLM2 / OLMo-3 checkpoints
80
46 / 40,000
21
Earlier control study
55
0 / 27,500
0
Table 6: Clipping diagnostics across experiments. A flagged step contains at least one selected token with a probability ratio outside [0.8, 1.2]. The count does not measure the resulting gradient change.
Checkpoint
Seed
Lenient
Strict
Format rate
P2 +5B
31003
67.5
0.0
100.0
P2 +5B
31004
59.0
0.0
0.0
P2 +5B
31006
77.5
0.0
0.0
P2 +5B
31007
77.0
0.0
50.5
P2 +9B
31004
78.0
43.0
54.0
P2 +9B
31008
55.0
4.5
9.5
Table 7: Answer-format diagnostics for all GSM Random runs ending at or above 50% lenient accuracy. Strict accuracy requires a standalone answer line; format rate records the presence of an answer marker. Values are percentages.
Checkpoint
Base
1
2
3
4
Max
n
P1 5B
0.0
0.0
0.0
0.0
0.5
0.5
4
P1 34B
0.5
2.0
0.5
0.0
1.0
2.0
4
P1 462B
3.5
4.5
3.5
4.5
8.5
8.5
4
P1 839B
3.5
55.0
33.5
16.5
19.5
55.0
4
P1 1,259B
2.0
66.5
72.5
1.5
57.0
72.5
4
P1 1,469B
2.5
18.5
18.0
12.0
19.5
19.5
4
Table 8: Synthetic GSM GT terminal accuracy at step 500 (%). Seed headings omit 3100; Max is the row maximum. Dashes mark unrun configurations. Base follows the frozen-score convention in Appendix 10.7.1 .
Checkpoint
1
2
3
4
5
6
7
8
Max
P1 5B
0.0
0.0
0.0
0.0
0.5
0.5
0.5
0.5
0.5
P1 34B
0.5
2.0
1.0
1.5
1.0
0.0
0.5
1.5
2.0
P1 462B
4.0
2.5
4.0
3.5
4.5
0.0
2.5
3.0
4.5
P1 839B
1.5
2.5
1.5
3.0
4.0
4.0
4.0
3.5
4.0
P1 1,259B
3.0
2.5
2.5
1.0
2.0
2.5
2.5
3.5
3.5
P1 1,469B
2.0
1.0
1.0
1.5
2.0
2.5
1.0
0.0
2.5
Table 9: Synthetic GSM Random terminal accuracy at step 500 (%). Seed headings omit 3100; Max is the row maximum. Dashes mark unrun configurations.
Figure 8 : Frozen and maximum terminal accuracy after 500 updates on code and lookup. Faint points show individual seeds. Lookup pools 104 one-hop and 96 two-hop items, with uniform choice among four cities shown at 25%. Panel scales differ. Checkpoints are equally spaced in training order; P2 tokens are additional to P1. Table 2 gives seed counts.
Checkpoint
Base
1
2
3
4
Max
n
P1 5B
0.0
0.0
—
—
—
0.0
1
P1 462B
0.5
0.5
—
—
—
0.5
1
P1 839B
0.0
0.0
—
—
—
0.0
1
P1 1,259B
0.5
0.5
—
—
—
0.5
1
P1 2,098B
1.0
11.0
—
—
—
11.0
1
P1 2,937B
0.0
0.5
—
—
—
0.5
1
Table 10: Code GT terminal accuracy at step 500 (%). Seed headings omit 3100; Max is the row maximum. Dashes mark unrun configurations. Base follows the frozen-score convention in Appendix 10.7.1 .
Checkpoint
1
2
3
4
5
6
7
8
Max
P1 5B
0.0
0.0
—
—
—
—
—
—
0.0
P1 462B
0.0
1.0
—
—
—
—
—
—
1.0
P1 839B
0.0
0.0
—
—
—
—
—
—
0.0
P1 1,259B
0.0
0.0
—
—
—
—
—
—
0.0
P1 2,098B
1.5
0.0
—
—
—
—
—
—
1.5
P1 2,937B
0.5
0.5
—
—
—
—
—
—
0.5
Table 11: Code Random terminal accuracy at step 500 (%). Seed headings omit 3100; Max is the row maximum. Dashes mark unrun configurations.
Checkpoint
Base
1
2
3
4
Max
n
P1 5B
0.0
0.0
0.0
—
—
0.0
2
P1 462B
3.0
97.5
97.0
—
—
97.5
2
P1 839B
22.5
98.5
99.5
—
—
99.5
2
P1 1,259B
27.0
99.0
98.0
—
—
99.0
2
P1 2,098B
20.0
99.5
99.0
—
—
99.5
2
P1 2,937B
5.0
99.0
98.5
—
—
99.0
2
Table 12: Compositional lookup GT terminal accuracy at step 500 (%). Seed headings omit 3100; Max is the row maximum. Dashes mark unrun configurations. Base follows the frozen-score convention in Appendix 10.7.1 .
Checkpoint
1
2
3
4
5
6
7
8
Max
P1 5B
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
P1 462B
26.0
2.0
25.0
1.0
3.0
5.5
5.5
4.0
26.0
P1 839B
20.5
20.5
22.0
0.0
27.0
23.5
22.0
1.5
27.0
P1 1,259B
17.0
24.5
3.5
33.0
17.5
33.0
24.0
0.0
33.0
P1 2,098B
12.0
11.5
16.5
19.0
15.0
15.5
10.0
12.5
19.0
P1 2,937B
0.0
2.0
3.0
3.0
23.0
16.0
7.0
0.0
23.0
Table 13: Compositional lookup Random terminal accuracy at step 500 (%). Seed headings omit 3100; Max is the row maximum. Dashes mark unrun configurations.
Figure 9 : Checkpoint development in SmolLM2 and OLMo-3. Curves show frozen greedy accuracy and maximum terminal accuracy after 500 updates; faint points show the two GT and eight Random seeds at each checkpoint. All use the same 200 GSM questions. Horizontal spacing follows checkpoint order; shading marks OLMo-3 mid-training.
Checkpoint
Raw
GT max
Random max
Random mean
Gain
Random peak
S1
0.5
1.0
1.0
0.56
0/8
1.0
S2
0.0
1.0
4.0
1.12
0/8
4.0
S3
1.0
3.0
8.5
2.44
0/8
8.5
S4
3.0
78.0
7.5
4.69
0/8
7.5
O1
1.0
0.5
1.0
0.63
0/8
1.0
O2
15.5
93.5
15.0
8.31
0/8
31.5
Table 14: Additional-model checkpoint results (%). Maxima use terminal scores, except Random peak, which also includes intermediate evaluations. Gain counts terminal improvements greater than ten points.
GT
Random
ID
1
2
1
2
3
4
5
6
7
8
S1
1.0
0.5
0.5
0.5
0.5
0.5
0.5
1.0
0.5
0.5
S2
0.5
1.0
0.5
0.5
2.0
1.5
0.0
0.0
0.5
4.0
S3
3.0
1.0
8.5
2.0
1.0
4.0
1.0
1.5
0.5
1.0
S4
54.0
78.0
7.0
3.5
2.5
5.0
7.5
5.0
5.0
2.0
O1
0.5
0.5
0.5
0.5
1.0
1.0
1.0
0.0
0.5
0.5
Table 15: All 80 terminal scores in the additional checkpoint comparison (%). GT columns use seeds 31001–31002 and Random columns use seeds 31001–31008.
Checkpoint
S1
S2
S3
S4
O1
O2
O3
O4
Full output
0.5
0.0
1.0
3.0
1.0
15.5
18.0
30.0
First block
3.0
4.0
14.0
27.5
1.0
53.0
35.0
34.0
Table 16: Frozen accuracy with the primary full-output extractor and the first-block diagnostic (%). The developmental figures retain the primary extractor before and after training.
Figure 10 : Terminal GSM accuracy under the five OLMo-2 input conditions. Dots show every seed; diamonds mark maxima. Original GSM and the two random-token conditions have 32 seeds each, and shuffled GSM and MMLU have 16. The dashed line is the frozen score of 20.5%.
Training input
n
Max
Mean
Median
Gain
Collapse
B: GSM
32
78.0
13.98
6.75
3/32
15/32
C: shuffled GSM
16
55.5
15.56
14.25
2/16
6/16
A: MMLU
16
86.5
23.28
18.50
2/16
1/16
D: tokens + math template
32
68.5
17.50
13.50
4/32
8/32
E: raw tokens
32
23.0
20.36
20.50
0/32
0/32
Table 17: OLMo-2 input-control outcomes at step 500. Scores are percentages. Gain means improvement greater than ten points from a run's initial score; collapse means terminal accuracy at or below 5%.
Seed
B
C
A
D
E
31001
23.5
8.5
23.0
16.5
19.5
31002
7.5
0.5
15.5
22.0
19.5
31003
15.0
21.0
22.5
8.5
23.0
31004
6.0
35.0
50.5
68.5
21.0
31005
0.0
20.5
19.0
23.0
20.5
31006
21.5
55.5
29.5
22.5
21.5
Table 18: All OLMo-2 input-control terminal scores (%). B uses original GSM, C shuffled GSM, A MMLU stems, D tokens within the math template, and E tokens without the template. Dashes indicate seeds not run.
Figure 11 : Training trajectories with unrelated MMLU questions and shuffled GSM questions. Both retain the math template. Thin lines show all 16 seeds; thick lines show the pointwise maximum, which can switch seeds. The dashed line marks frozen accuracy.
Figure 12 : Random-token input controls on Qwen2.5 and Llama. Dots show the eight terminal scores in each condition; diamonds mark maxima. Dashed lines show mean pre-update scores, 44.3% and 25.5%. Input construction differs in both template and token filtering (Appendix 10.3 ).
Model
Input
Max
Mean
SD
Gain
Collapse
Qwen2.5-7B
D
91.5
50.75
29.78
4/8
1/8
E
51.0
46.38
2.45
0/8
0/8
Llama-3.1-8B
D
64.0
16.63
19.30
1/8
3/8
E
31.5
19.56
6.94
0/8
0/8
Table 19: Qwen2.5 and Llama input-control results (%), with eight seeds per condition. D retains the math template; E omits it. Gain means an improvement above ten points; collapse means terminal accuracy at or below 5%.
Qwen2.5-7B
Llama-3.1-8B
Seed
D
E
D
E
31001
57.0
51.0
1.0
31.5
31002
34.0
43.0
3.0
14.5
31003
0.0
45.0
1.0
9.0
31004
25.5
48.0
19.0
16.0
31005
91.5
44.0
13.5
18.0
Table 20: All Qwen2.5 and Llama input-control terminal scores (%). D uses tokens within the math template and E uses tokens alone.
Source
Paired quantity (P2 minus P1)
Counts
p
Holm p
P1
Gain
6/9/1
0.6072
0.8405
P1
Improved
3/0
0.2500
0.8405
P1
Absolute gain
13/2/1
0.0074
0.0443
P2
Gain
11/5/0
0.2101
0.8405
P2
Improved
5/1
0.2188
0.8405
P2
Absolute gain
12/4/0
0.0768
0.3841
Table 21: The six planned shared-rollout tests, with 16 paired blocks per source. Sign-test counts are positive/negative/tied differences. Improved-run counts are pairs where only P2 improves / only P1 improves. Holm correction covers all six tests.
Source
Learner
Primary score
Prefix score
Prefix gain
Prefix gain
before → after
before → after
>10 pp
<−10 pp
P1
P1
2.5→3.1
41.5→28.8
2/16
10/16
P1
P2
13.5→18.0
14.5→38.4
8/16
0/16
P2
P1
2.5→5.4
41.5→21.2
1/16
11/16
P2
P2
13.5→27.2
14.5→39.3
8/16
0/16
Table 22: Shared-rollout answer-extraction diagnostic. Scores are percentages; the terminal score is the mean of 16 runs. Prefix scoring stops at the first generated newline followed by Question: . These descriptive results use the same rule before and after training.
Replicate
P1 → P1
P1 → P2
P2 → P1
P2 → P2
51001
3.0
4.0
1.5
48.5
51002
4.0
44.0
7.0
14.5
51003
4.0
23.0
1.0
13.5
51004
8.0
2.0
4.0
10.5
51005
1.0
43.5
3.0
47.5
51006
0.5
11.5
7.0
19.0
Table 23: All shared-rollout terminal scores (%) on the 200 held-out GSM questions. Arrows give response source → learner. Each row is one paired block of four runs.
seed
P1 (3,896B)
P2 +5B
P2 +50B
r_g4 ( G = 4)
777
3.0 / 7.5 (+4.5)
13.0 / 26.5 (+13.5) ⋆
21.0 / 1.5 ( − 19.5)
888
1.0 / 27.0 (+26.0) ⋆
9.0 / 2.5 ( − 6.5)
17.5 / 0.5 ( − 17.0)
999
4.5 / 9.0 (+4.5)
20.0 / 17.5 ( − 2.5)
27.0 / 76.5 (+49.5) ⋆
1001
4.5 / 0.0 ( − 4.5)
14.0 / 19.5 (+5.5)
19.5 / 6.0 ( − 13.5)
1002
1.0 / 8.0 (+7.0)
14.5 / 49.0 (+34.5) ⋆
16.0 / 85.5 (+69.5) ⋆
Table 24: All 30 random-reward runs in the earlier group-size comparison. Entries give initial / terminal accuracy in percent, with changes in parentheses. Stars mark gains above 10 points. Every run uses learning rate 3×10−5 , reference coefficient 0.01, and 500 updates.
Checkpoint
G = 4 successes
G = 4 mean Δ
G = 8 successes
G = 8 mean Δ
P1 3,896B
1/5
+7.5
0/5
+0.8
P2 +5B
2/5
+8.9
0/5
− 4.0
P2 +50B
2/5
+13.8
0/5
− 6.1
Table 25: Group-size outcomes at each checkpoint, with five seeds per condition. Success means a terminal gain above 10 points; mean changes are in percentage points.
arm
after accuracy (777 / 888 / 999 / 1001 / 1002)
mean Δ
gt_g4 P1
90.5 / 92.5 / 81.5 / 81.0 / 84.0
+83.1
gt_g4 P2 +5B
99.0 / 97.0 / 98.0 / 99.5 / 100.0
+84.6
gt_g4 P2 +50B
100.0 / 98.0 / 100.0 / 98.0 / 21.5
+63.3
gt_K50
90.5 / 1.0 / 62.0 / 3.0 / 54.0
+28.0
anti_K50
1.5 / 1.5 / 16.5 / 3.5 / 9.0
− 7.7
Table 26: GT and early-steering outcomes in the earlier study. Terminal accuracies follow seed order 777, 888, 999, 1001, 1002 and are reported in percent; mean changes are in percentage points. Steering is applied at P2 +5B.
Quantity
D
E
Difference
p
pH
pall
OLMo-2
Mean signed gain (pp)
− 3.00
− 0.14
− 2.86
0.3337
1
1
Mean absolute change (pp)
12.59
1.11
+11.48
4.66×10−10
6.98×10−9
1.35×10−7
Gain >10 pp
4/32
0/32
+12.50
0.125
1
1
Absolute change >10 pp
16/32
0/32
+50.00
3.05×10−5
4.27×10−4
0.0088
Terminal score at most 5%
8/32
0/32
+25.00
0.0078
0.1016
1
Table 27: Paired D/E comparisons on OLMo-2 (32 seeds), Qwen2.5 (eight), and Llama (eight). Effects are D minus E in percentage points; binary rows also report event counts. pH corrects the 15 tests in this table, and pall corrects all 290 exploratory tests.
Model
Input
Gain: count [95% CI]
Change: count [95% CI]
OLMo-2
B
3/32 [2.0, 25.0]
21/32 [46.8, 81.4]
OLMo-2
C
2/16 [1.6, 38.3]
10/16 [35.4, 84.8]
OLMo-2
A
2/16 [1.6, 38.3]
5/16 [11.0, 58.7]
OLMo-2
D
4/32 [3.5, 29.0]
16/32 [31.9, 68.1]
OLMo-2
E
0/32 [0.0, 10.9]
0/32 [0.0, 10.9]
Qwen2.5
D
4/8 [15.7, 84.3]
7/8 [47.3, 99.7]
Table 28: Input-control event counts and marginal 95% Clopper–Pearson intervals (%). Gain means Δ>10 points; change means ∣Δ∣>10 points. Each interval describes one model, input condition, and protocol.
Pair
n
Mean diff.
p
pH
Rate diff.
p
pH
B − C
16
− 2.59
0.5857
1
− 6.25
1
1
B − A
16
− 10.31
0.0023
0.0417
− 6.25
1
1
B − D
32
− 3.64
0.4467
1
− 3.12
1
1
B − E
32
− 6.50
0.0672
1
+9.38
0.25
1
C − A
16
− 7.72
0.1971
1
0.00
1
1
C − D
16
− 5.16
0.3455
1
0.00
1
1
Table 29: Other paired OLMo-2 input comparisons, using shared seeds. The first triplet gives mean signed-gain differences; the second gives differences in the rate of gains above ten points. Effects are percentage points. pH corrects all 18 tests; D/E appears separately in Table 27 .
Comparison
Seeds
Difference
p
pH
GSM GT: P1 3,896B − P1 34B
4
+81.12
0.125
0.875
GSM GT: P2 +5B − P1 3,896B
4
+17.38
0.125
0.875
GSM GT: P1 839B − P1 462B
4
+25.88
0.125
0.875
GSM Random: P1 3,896B − P1 34B
8
+5.81
0.0156
0.1719
GSM Random: P2 +5B − P1 3,896B
8
+33.44
0.0156
0.1719
GSM Random: P1 839B − P1 462B
8
0.00
1
1
Table 30: OLMo-2 developmental contrasts. Effects are later minus earlier terminal scores in percentage points, except the last row, which compares high-score frequencies. P1/P2 contrasts average checkpoints within phase and seed before pairing. pH corrects 12 tests. Code GT has one common seed and no inferential test.
Pair
GT diff.
p
pH
Random diff.
p
pH
S2 − S1
0.00
1
1
+0.56
0.4375
1
S3 − S1
+1.25
0.5
1
+1.88
0.0156
0.4219
S4 − S1
+65.25
0.5
1
+4.12
0.0078
0.25
S3 − S2
+1.25
1
1
+1.31
0.3594
1
S4 − S2
+65.25
0.5
1
+3.56
0.0234
0.5859
S4 − S3
+64.00
0.5
1
+2.25
0.0469
1
Table 31: All checkpoint pairs within SmolLM2 and OLMo-3. Triplets give mean terminal-score differences (percentage points), raw p , and Holm p for GT and Random. GT uses two paired seeds and Random eight. Correction covers these 24 tests and the eight GT/Random tests in Table 36 .
Comparison
Mean diff.
p
pH
Rate diff.
p
pH
P1 3,896B
+6.70
0.3125
1
+20.00
1
1
P2 +5B
+12.90
0.1875
1
+40.00
0.5
1
P2 +50B
+19.90
0.4375
1
+40.00
0.5
1
Three-checkpoint mean
+13.17
0.3125
1
+33.33
0.125
1
GT − anti steering
+35.70
0.25
1
+60.00
0.25
1
Table 32: Group-size and steering comparisons with five seed blocks. The first four rows compare G = 4 with G = 8, including an average over checkpoints; the last compares GT and inverted-GT steering. Triplets report mean-gain and jackpot-rate differences (percentage points), raw p , and Holm p over all ten tests.
Checkpoint
GT ↑/↓/=
pdir
pH
R ↑/↓/=
pdir
pH
GT − R
p
pH
P1 5B
1/0/3
1
1
4/0/4
0.125
1
+0.12
1
1
P1 34B
2/1/1
1
1
5/1/2
0.2188
1
− 0.38
0.75
1
P1 462B
3/0/1
0.25
1
3/4/1
1
1
+1.75
0.125
1
P1 839B
4/0/0
0.125
1
3/4/1
1
1
+29.00
0.125
1
P1 1,259B
3/1/0
0.625
1
6/1/1
0.125
1
+47.12
0.25
1
P1 1,469B
4/0/0
0.125
1
1/5/2
0.2188
1
+15.62
0.125
1
Table 33: GSM changes within each checkpoint and reward condition. Tables 34 – 36 use the same layout. Up/down/tie counts use each run's initial score. pdir is a two-sided sign test; its pH corrects 150 within-condition tests. The final triplet gives the paired GT minus Random mean score (percentage points), permutation p , and correction over 53 scan GT/Random tests. One-seed comparisons have no p value.
Checkpoint
GT ↑/↓/=
pdir
pH
R ↑/↓/=
pdir
pH
GT − R
p
pH
P1 5B
0/0/1
–
–
0/0/2
1
1
0.00
–
–
P1 462B
0/0/1
–
–
1/1/0
1
1
+0.50
–
–
P1 839B
0/0/1
–
–
0/0/2
1
1
0.00
–
–
P1 1,259B
0/0/1
–
–
0/2/0
0.5
1
+0.50
–
–
P1 2,098B
1/0/0
–
–
1/1/0
1
1
+9.50
–
–
P1 2,937B
1/0/0
–
–
2/0/0
0.5
1
0.00
–
–
Table 34: Code changes within each checkpoint and reward condition. Columns and tests follow Table 33 .
Checkpoint
GT ↑/↓/=
pdir
pH
R ↑/↓/=
pdir
pH
GT − R
p
pH
P1 5B
0/0/2
1
1
0/0/8
1
1
0.00
1
1
P1 462B
2/0/0
0.5
1
5/3/0
0.7266
1
+83.25
0.5
1
P1 839B
2/0/0
0.5
1
2/5/1
0.4531
1
+78.50
0.5
1
P1 1,259B
2/0/0
0.5
1
2/6/0
0.2891
1
+77.75
0.5
1
P1 2,098B
2/0/0
0.5
1
0/8/0
0.0078
1
+87.50
0.5
1
P1 2,937B
2/0/0
0.5
1
3/5/0
0.7266
1
+97.75
0.5
1
Table 35: Lookup changes within each checkpoint and reward condition. Columns and tests follow Table 33 .
Checkpoint
GT ↑/↓/=
pdir
pH
R ↑/↓/=
pdir
pH
GT − R
p
pH
O1
0/2/0
0.5
1
0/5/3
0.0625
1
0.00
1
1
O2
2/0/0
0.5
1
0/8/0
0.0078
1
+79.25
0.5
1
O3
2/0/0
0.5
1
3/5/0
0.7266
1
+49.25
0.5
1
O4
2/0/0
0.5
1
4/4/0
1
1
+44.50
0.5
1
S1
1/0/1
1
1
1/0/7
1
1
+0.25
1
1
S2
2/0/0
0.5
1
6/0/2
0.0312
1
+0.25
1
1
Table 36: SmolLM2 and OLMo-3 changes within each checkpoint and reward condition. Columns follow Table 33 , except that the GT minus Random correction covers 32 additional-model tests.
Cell
Mean gain
↑/↓/=
pdir
pH
Gain: count [95% CI]
OLMo-2 B
− 6.64
6/25/1
8.78×10−4
0.1317
3/32 [2.0, 25.0]
OLMo-2 C
− 4.94
6/9/1
0.6072
1
2/16 [1.6, 38.3]
OLMo-2 A
+2.78
7/9/0
0.8036
1
2/16 [1.6, 38.3]
OLMo-2 D
− 3.00
13/19/0
0.3771
1
4/32 [3.5, 29.0]
OLMo-2 E
− 0.14
13/13/6
1
1
0/32 [0.0, 10.9]
Qwen2.5 D
+6.56
4/4/0
1
1
4/8 [15.7, 84.3]
Table 37: Changes within each input-control and earlier-study condition. Mean gains are percentage points. Sign tests compare upward with downward nonzero changes, with Holm correction over 150 tests. The final column gives the jackpot count and its marginal 95% interval (%).
Checkpoint / task
Gain: count [95% CI]
High score: count [95% CI]
GSM P1 5B
0/8 [0.0, 36.9]
0/8 [0.0, 36.9]
GSM P1 34B
0/8 [0.0, 36.9]
0/8 [0.0, 36.9]
GSM P1 462B
0/8 [0.0, 36.9]
0/8 [0.0, 36.9]
GSM P1 839B
0/8 [0.0, 36.9]
0/8 [0.0, 36.9]
GSM P1 1,259B
0/8 [0.0, 36.9]
0/8 [0.0, 36.9]
GSM P1 1,469B
0/8 [0.0, 36.9]
0/8 [0.0, 36.9]
Table 38: Random-run success rates, part 1. Entries give success count / seed count and marginal 95% Clopper–Pearson intervals (%). Gain means Δ>10 points and high score means terminal accuracy 50% or more. Each checkpoint is estimated separately.
Checkpoint / task
Gain: count [95% CI]
High score: count [95% CI]
GSM P1 3,896B
0/8 [0.0, 36.9]
0/8 [0.0, 36.9]
GSM P2 +5B
4/8 [15.7, 84.3]
4/8 [15.7, 84.3]
GSM P2 +9B
2/8 [3.2, 65.1]
2/8 [3.2, 65.1]
GSM P2 +13B
0/8 [0.0, 36.9]
0/8 [0.0, 36.9]
GSM P2 +17B
0/8 [0.0, 36.9]
0/8 [0.0, 36.9]
GSM P2 +21B
3/8 [8.5, 75.5]
3/8 [8.5, 75.5]
Table 39: Random-run success rates, part 2. Entries give success count / seed count and marginal 95% Clopper–Pearson intervals (%). Gain means Δ>10 points and high score means terminal accuracy 50% or more. Each checkpoint is estimated separately.
Checkpoint / task
Gain: count [95% CI]
High score: count [95% CI]
CODE P2 +50B
1/8 [0.3, 52.7]
0/8 [0.0, 36.9]
Lookup P1 5B
0/8 [0.0, 36.9]
0/8 [0.0, 36.9]
Lookup P1 462B
2/8 [3.2, 65.1]
0/8 [0.0, 36.9]
Lookup P1 839B
0/8 [0.0, 36.9]
0/8 [0.0, 36.9]
Lookup P1 1,259B
0/8 [0.0, 36.9]
0/8 [0.0, 36.9]
Lookup P1 2,098B
0/8 [0.0, 36.9]
0/8 [0.0, 36.9]
Table 40: Random-run success rates, part 3. Entries give success count / seed count and marginal 95% Clopper–Pearson intervals (%). Gain means Δ>10 points and high score means terminal accuracy 50% or more. Each checkpoint is estimated separately.
Sparse reward reinforcement learning (RL) has become a standard tool for improving LLM reasoning, but its success depends critically on the coverage present in the base model. In practice, models are often primed for RL through \emph{mid-training} on curated reasoning traces that teach useful primitive skills such as decomposition, verification, or self-correction. Although effective, this strategy requires manually specifying what the model should learn, and it remains unclear whether such primitive coverage is enough for much harder problems, which require combining these skills into broader solution strategies. We study a more automated approach: \emph{RL-based mid-training} using large corpora of human-written question-answer data. Rather than treating reference solutions as targets to imitate, our method, ExpRL, uses them as \emph{reward scaffolds}: references are hidden from the policy and used only to construct problem-specific grading rubrics for judging on-policy reasoning traces. The policy samples from the original problem prompt, while an LLM judge compares the sampled reasoning trace against the reference solution and assigns outcome-level or process-level dense rewards. This lets ExpRL reinforce partial progress, useful intermediate reductions, and productive reasoning behaviors that sparse final-answer rewards often fail to upweight. On challenging math reasoning tasks, ExpRL yields stronger RL priming than SFT, sparse-reward GRPO, and self-distillation, and provides a better initialization for subsequent sparse-reward RL. Additional mixed-domain experiments further suggest that ExpRL can extend beyond the original math-only setting.
Violet Xiang, Amrith Setlur, Chase Blagden +2
1Stanford University · 2Carnegie Mellon University · 3OpenAI +1
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, but prompt groups with identical rollout rewards consume generation budget without effective learning signals. Pre-rollout prompt selection can reduce this waste by screening prompts before rollout generation. However, existing pre-rollout methods struggle to balance exploitation and exploration: repeatedly exploiting historically informative prompts can narrow training coverage, whereas broader exploration can lower the fraction of informative prompts. To address these limitations, we introduce LEEPS, a Latent-Guided Explore--Exploit Prompt Sampler that adaptively balances the reuse of previously observed informative prompts with continued exploration of uncertain ones. LEEPS partitions candidates into exploit and explore portfolios and adaptively allocates rollout budget according to their recent non-trivial ratios. It further uses representation-space neighbors and historical rollout outcomes to prioritize uncertain prompts likely to yield non-zero reward variance, thereby making exploration more targeted without additional rollouts. Across six mathematical reasoning benchmarks, LEEPS achieves the highest average score at both model scales, with relative gains of 2.6% and 3.7% over the strongest baseline for Qwen2.5-Math-1.5B and 7B, respectively, and generally improves faster during the training process. It also achieves the highest average score across the three evaluated OOD general-reasoning benchmarks at both model scales and adds only about 2 seconds of online sampling overhead per training step. Code is available at https://github.com/ShuangLiangX/LEEPS.
Reinforcement learning with verifiable rewards, particularly Group Relative Policy Optimization (GRPO), has significantly advanced the reasoning capabilities of Large Language Models (LLMs). However, in complex tasks, GRPO frequently suffers from the ``zero-advantage problem'': when all sampled rollouts for a query fail, the relative advantage collapses to zero. Consequently, the model loses effective training signals for these questions, wasting the training data and computational budget. While simply increasing the sampling budget for these questions is a common remedy, the static sampling policy inherently constrains reasoning exploration, limiting the success rate. In this paper, we propose Lorem Perturbation for Exploration (LoPE), a simple yet effective training framework to break this exploration bottleneck. We posit that task-irrelevant prompt-space perturbations can shift the model's output distribution enough to unlock orthogonal reasoning pathways for hard questions. Specifically, LoPE prepends sequences stochastically assembled from Lorem Ipsum vocabulary (a pseudo-Latin placeholder text) to the prompts before resampling. Experiments across 1.7B, 4B, and 7B models demonstrate that LoPE significantly outperforms resampling with the original prompts. Further analysis reveals that other Latin-based random sequences with low perplexity are also effective perturbations. Our results establish LoPE as a strong baseline for broadening exploration in LLM reinforcement learning.