Reinforcement Learning with Verifiable Rewards (RLVR) performs well on problems with clear rewards, such as mathematics and coding, but whether it also works where the reward is less clear remains open. The reason-over-search recipe applies RLVR to open-domain question answering, where retrieval grounds the answer and a match against the reference supplies the reward. So far it has been demonstrated on large models, and below one billion parameters only with distillation from a larger teacher. We test the recipe on a small model. We train Qwen3.5-0.8B with Group Relative Policy Optimization (GRPO) and an interleaved Wikipedia-search tool on MuSiQue, varying only the reward across three shapes over three seeds each, and we evaluate every checkpoint held-out on a seven-benchmark question-answering suite. The recipe works: the best run reaches 0.352 average exact match against a 0.092 untrained floor, a 3.8-fold gain, with no distillation step in the training loop. The reward shape also matters. The Search-R1-faithful exact-match-only reward is the worst of the three at every seed at the matched training horizon, and it is worst even on exact match, the metric it directly optimises. We conclude that the sparse exact-match reward, RLVR's default in mathematics and code, is the wrong starting point for models of this size. The reason-over-search setting can supply a suitable reward for RLVR on small models, but small-model RLVR needs its own reward-design study rather than a scaled-down copy of a large-model recipe.
Figures & tables
Figure 1: The critic-free GRPO loop. A group of G=5 rollouts per question is scored; the reward node branches into the three ablated shapes (F1 with a 0.1 format floor, F1-only, and sparse EM-only); the group mean and standard deviation set the advantage baseline, with no value critic; the update is the clipped, KL-regularised surrogate of Section 3 . A worked reason-search-read rollout, showing the interleaved tool calls and the loss masking of retrieved spans, is in Appendix 0.G .
Reward shape
seed 42
seed 43
seed 44
mean
range
Matched horizon (step ≤180 )
F1+format
0.322
0.292
0.308
0.307
0.030
F1-only
0.313
0.283
0.318
0.305
0.035
EM-only
0.301
0.264
0.249
0.271
0.052
Each seed’s own horizon
F1-only
0.313
0.291
0.352
0.318
0.061
Table 1: Held-out average EM by reward shape under both horizon rules. Upper panel: the matched horizon (best checkpoint at step ≤180 ), where EM-only is worst at every seed. Lower panel: each seed’s own horizon (seed 42 at ≤180 , 43 at ≤230 , 44 at ≤312 ), where F1-only owns the single best run and F1+format has the tightest spread.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Seed 42 held-out EM by reward shape (to step 180); F1+format leads and all three shapes clear the 0.092 floor early.
Figure 3: Seed 43 held-out EM by reward shape (to step 230); F1+format leads and is the most stable, while F1-only is the noisiest leg, with a deep collapse over steps 130 to 190.
Figure 4: Seed 44 held-out EM by reward shape (full epoch, to step 310); F1-only reaches the run-high 0.352 at step 310, and EM-only rises late, on the longest leg in the study.
Figure 5: Cross-seed mean with minimum-to-maximum range per reward shape, matched to step ≤180 , for reward, tool calls, and response length. The F1 shapes lean while reward climbs; EM-only stays highest on tool calls (reward levels are not comparable across shapes, Section 6.5 ).
Figure 6: Seed 42 training dynamics, capped at step 180.
Figure 7: Seed 43 training dynamics, capped at step 230.
Figure 8: Seed 44 training dynamics (full epoch; last full cadence at step 310). Reward climbs while tool calls, length, total tokens, and clip ratio all fall for the F1 shapes, while EM-only stays high.
Quantity (step 1 → 310)
F1-only
F1+format
EM-only
Tool calls per sample
6.6 → 3.4
6.5 → 4.3
6.5 → 6.2
Total tokens per sample
6435 → 3694
6360 → 4730
6380 → 6031
Length-clip ratio
0.63 → 0.01
0.63 → 0.01
0.62 → 0.36
No-valid-answer ratio
0.72 → 0.01
0.72 → 0.01
0.72 → 0.37
Appendix
Table 2: Seed-44 cold-start-to-end magnitudes (step 1 to step 310). The F1 shapes compress on every axis shown; EM-only leans mid-horizon and then re-enters over-search. Per-rollout response length is not shown because it does not move in one direction across shapes (Section 6.4 ).
Figure 9: Matched-horizon average EM by reward shape, one bar per seed (the same cells as the upper panel of Table 1 ). EM-only is the shortest group at every seed; the bar spread shows seed variance (F1+format tightest, EM-only widest). All cells are best checkpoints under the same step ≤180 cap, so this is not a training-length effect.
Figure 10: Matched-horizon cross-seed mean of EM, token-F1, and substring accuracy at each run’s best-EM checkpoint, by reward shape; dashed lines mark each metric’s untrained floor. EM-only is lowest on all three metrics and the two F1 shapes are close and higher, so the ranking is not an artifact of scoring on the metric EM-only optimises.
Figure 11: Seed-42 training-health diagnostics by reward shape: gradient norm, loss, KL divergence, approximate policy entropy, aborted ratio, and natural-termination rate. Seed 42 is shown; the bounded-gradient-norm and lowest-EM-only-gradient-norm reads are computed over all nine runs (the other seeds, not shown, are similar).
Figure 12: One reason-search-read rollout, with the retrieved spans shaded to mark the loss mask. The reward is computed on the final answer block only. This panel and Figure 1 are two halves of one source drawing, so its in-figure lettering is inherited: “panel a” is Figure 1 , and the “Appendix C” it cites for the piecewise reward definition is a stale pointer, superseded here by Equation 4 .
Figure 13: Held-out behaviour by search count at each run’s best-EM checkpoint, pooled over three seeds and the four multi-hop benchmarks. Left: mean exact match against search count for completed rollouts; all three shapes peak at two searches. Right: the search-count distribution per shape with the per-shape abort rate; EM-only places far more mass in the aborted turn-budget bin, the main source of its lower held-out exact match.
Run
NQ
TriQA
PopQA
HotQA
2Wiki
MuSiQue
Bambgl
avg EM
F1-only s42
0.301
0.492
0.319
0.309
0.287
0.130
0.352
0.313 @110
F1-only s43
0.301
0.489
0.321
0.259
0.274
0.093
0.296
0.291 @210
F1-only s44
0.363
0.531
0.370
0.359
0.285
0.161
0.392
0.352 @310
F1+format s42
0.319
0.500
0.335
0.311
0.313
0.117
0.360
0.322 @100
F1+format s43
0.335
0.533
0.352
0.322
0.296
0.092
0.264
0.313 @210
F1+format s44
0.310
0.521
0.340
0.307
0.278
0.121
0.288
0.309 @270
Appendix
Table 3: Per-benchmark EM at each run’s seed-horizon best checkpoint (step in the avg-EM column). MuSiQue is the single in-distribution set; the other six are out-of-distribution, so the generalisation claim rests on them.
Reward shape
seed 42
seed 43
seed 44
mean
range
F1-only
0.402
0.369
0.407
0.393
0.038
F1+format
0.413
0.370
0.390
0.391
0.043
EM-only
0.368
0.333
0.319
0.340
0.049
Appendix
Table 4: Matched horizon, average token-F1 per reward shape at each run’s best-F1 checkpoint, showing the same ordering on the dense metric.
Reward shape
seed 42
seed 43
seed 44
mean
F1-only
0.394
0.355
0.397
0.382
F1+format
0.378
0.359
0.355
0.364
EM-only
0.334
0.293
0.340
0.323
Appendix
Table 5: Substring-cover accuracy (ACC) at each run’s seed-horizon best-EM checkpoint. EM-only is worst on the most lenient metric as well.
Reinforcement learning with verifiable rewards (RLVR) optimizes LLMs using sparse verifiable final-answer rewards. This sparse anchor reliably verifies whether a trajectory succeeds but provides no direct feedback on the reasoning path that produced it. Before success, prerequisite progress on hard problems receives no reward signal; after success, outcome rewards cannot distinguish well-organized correct trajectories from redundant or locally flawed ones. We introduce SCOPE-RL (Scaffolded Chain Optimization with Process Efficiency), a two-stage framework that densifies this anchor while retaining the GRPO update: Adaptive Scaffolded RL adds prefix-decomposed verifiable rewards on answer-hidden sub-question chains before success, and Quality-Aware Process RL applies correctness-gated process-shape rewards to refine correct trajectories after success. An expert-validated Step-Quality Evaluation Protocol evaluates useful-step density, error localization, and token efficiency beyond final-answer accuracy. On Qwen3-8B-Instruct trained on DAPO-Math and Big-Math, SCOPE-RL improves average accuracy by up to 11.2 pp and reduces reasoning tokens by up to 27.1% over outcome-only GRPO; the gains hold under GSPO and on Qwen3-0.6B-Instruct, indicating that reward-signal densification is complementary to policy-update-level RLVR advances. Code and data are available at https://github.com/tokencraft-lab/SCOPE-RL.
Reinforcement Learning with Verifiable Rewards (RLVR) is widely used for post-training Large Language Models, but practical verifiers can make errors. We study how the rate and structure of reward noise affect RLVR in code generation, with a preliminary scientific-reasoning check. In multi-seed Qwen3 8B experiments on MBPP, mean validation reward over two fixed late evaluations is within 1 percentage point of the clean baseline at the tested resampled group-rollout noise rates through 20%, and within about 2 points at 30%. Confidence intervals allow larger losses; these point estimates do not establish a general tolerance threshold. We also examine full-program pass@k, four controlled noise structures, two model-based verifiers, and policy models from three families spanning 4B-9B parameters. We derive conditional advantage distributions for symmetric and asymmetric group noise, including retained format penalties, and show why clipping limits simple gradient-scaling arguments. The analysis identifies information preserved by whole-group corruption and limits on interpreting our asymmetric sweep as a precision-recall comparison. Overall, the results indicate that imperfect verification can support effective RLVR in the tested settings, while aggregate error rates alone do not characterize the learning signal.
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for improving mathematical reasoning in language models. Yet most RLVR work rewards only the final answer (outcome-based rewards), leaving the impact of step-level process supervision (process rewards) underexplored especially for small models that lack the capacity to self-correct under sparse feedback. We systematically compare five reward conditions applied to Qwen2.5-0.5B fine-tuned with Group Relative Policy Optimization (GRPO) on GSM8K: a no-RL baseline, process-only, outcome-only, and three hybrid weightings (λ∈{0.9,0.5,0.1} process weight). Process-only supervision achieves 63.73% test accuracy versus 53.75% for outcome-only, a nearly 10-percentage point gap while yielding reasoning traces with higher step validity and lower deviation from ground-truth chain length. Hybrid rewards generally correlate positively with process weight, with one notable anomaly: the low-process / high-outcome configuration (λ=0.1) underperforms pure outcome supervision, suggesting conflicting optimization signals. Error analysis using GPT-4o as a judge reveals distinct failure mode distributions: process models generate structurally inconsistent but arithmetically grounded traces, while outcome models produce concise but derivation-error-prone chains. Our results demonstrate that reward granularity is a first-order design decision for RLVR, with process-level supervision substantially improving both accuracy and trace fidelity in small language models.