Reinforcement Learning with Verifiable Rewards (RLVR) performs well on problems with clear rewards, such as mathematics and coding, but whether it also works where the reward is less clear remains open. The reason-over-search recipe applies RLVR to open-domain question answering, where retrieval grounds the answer and a match against the reference supplies the reward. So far it has been demonstrated on large models, and below one billion parameters only with distillation from a larger teacher. We test the recipe on a small model. We train Qwen3.5-0.8B with Group Relative Policy Optimization (GRPO) and an interleaved Wikipedia-search tool on MuSiQue, varying only the reward across three shapes over three seeds each, and we evaluate every checkpoint held-out on a seven-benchmark question-answering suite. The recipe works: the best run reaches 0.352 average exact match against a 0.092 untrained floor, a 3.8-fold gain, with no distillation step in the training loop. The reward shape also matters. The Search-R1-faithful exact-match-only reward is the worst of the three at every seed at the matched training horizon, and it is worst even on exact match, the metric it directly optimises. We conclude that the sparse exact-match reward, RLVR's default in mathematics and code, is the wrong starting point for models of this size. The reason-over-search setting can supply a suitable reward for RLVR on small models, but small-model RLVR needs its own reward-design study rather than a scaled-down copy of a large-model recipe.
Figures & tables
Figure 1: The critic-free GRPO loop. A group of G=5 rollouts per question is scored; the reward node branches into the three ablated shapes (F1 with a 0.1 format floor, F1-only, and sparse EM-only); the group mean and standard deviation set the advantage baseline, with no value critic; the update is the clipped, KL-regularised surrogate of Section 3 . A worked reason-search-read rollout, showing the interleaved tool calls and the loss masking of retrieved spans, is in Appendix 0.G .
Reward shape
seed 42
seed 43
seed 44
mean
range
Matched horizon (step ≤180 )
F1+format
0.322
0.292
0.308
0.307
0.030
F1-only
0.313
0.283
0.318
0.305
0.035
EM-only
0.301
0.264
0.249
0.271
0.052
Each seed’s own horizon
F1-only
0.313
0.291
0.352
0.318
0.061
Table 1: Held-out average EM by reward shape under both horizon rules. Upper panel: the matched horizon (best checkpoint at step ≤180 ), where EM-only is worst at every seed. Lower panel: each seed’s own horizon (seed 42 at ≤180 , 43 at ≤230 , 44 at ≤312 ), where F1-only owns the single best run and F1+format has the tightest spread.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Seed 42 held-out EM by reward shape (to step 180); F1+format leads and all three shapes clear the 0.092 floor early.
Figure 3: Seed 43 held-out EM by reward shape (to step 230); F1+format leads and is the most stable, while F1-only is the noisiest leg, with a deep collapse over steps 130 to 190.
Figure 4: Seed 44 held-out EM by reward shape (full epoch, to step 310); F1-only reaches the run-high 0.352 at step 310, and EM-only rises late, on the longest leg in the study.
Figure 5: Cross-seed mean with minimum-to-maximum range per reward shape, matched to step ≤180 , for reward, tool calls, and response length. The F1 shapes lean while reward climbs; EM-only stays highest on tool calls (reward levels are not comparable across shapes, Section 6.5 ).
Figure 6: Seed 42 training dynamics, capped at step 180.
Figure 7: Seed 43 training dynamics, capped at step 230.
Figure 8: Seed 44 training dynamics (full epoch; last full cadence at step 310). Reward climbs while tool calls, length, total tokens, and clip ratio all fall for the F1 shapes, while EM-only stays high.
Quantity (step 1 → 310)
F1-only
F1+format
EM-only
Tool calls per sample
6.6 → 3.4
6.5 → 4.3
6.5 → 6.2
Total tokens per sample
6435 → 3694
6360 → 4730
6380 → 6031
Length-clip ratio
0.63 → 0.01
0.63 → 0.01
0.62 → 0.36
No-valid-answer ratio
0.72 → 0.01
0.72 → 0.01
0.72 → 0.37
Appendix
Table 2: Seed-44 cold-start-to-end magnitudes (step 1 to step 310). The F1 shapes compress on every axis shown; EM-only leans mid-horizon and then re-enters over-search. Per-rollout response length is not shown because it does not move in one direction across shapes (Section 6.4 ).
Figure 9: Matched-horizon average EM by reward shape, one bar per seed (the same cells as the upper panel of Table 1 ). EM-only is the shortest group at every seed; the bar spread shows seed variance (F1+format tightest, EM-only widest). All cells are best checkpoints under the same step ≤180 cap, so this is not a training-length effect.
Figure 10: Matched-horizon cross-seed mean of EM, token-F1, and substring accuracy at each run’s best-EM checkpoint, by reward shape; dashed lines mark each metric’s untrained floor. EM-only is lowest on all three metrics and the two F1 shapes are close and higher, so the ranking is not an artifact of scoring on the metric EM-only optimises.
Figure 11: Seed-42 training-health diagnostics by reward shape: gradient norm, loss, KL divergence, approximate policy entropy, aborted ratio, and natural-termination rate. Seed 42 is shown; the bounded-gradient-norm and lowest-EM-only-gradient-norm reads are computed over all nine runs (the other seeds, not shown, are similar).
Figure 12: One reason-search-read rollout, with the retrieved spans shaded to mark the loss mask. The reward is computed on the final answer block only. This panel and Figure 1 are two halves of one source drawing, so its in-figure lettering is inherited: “panel a” is Figure 1 , and the “Appendix C” it cites for the piecewise reward definition is a stale pointer, superseded here by Equation 4 .
Figure 13: Held-out behaviour by search count at each run’s best-EM checkpoint, pooled over three seeds and the four multi-hop benchmarks. Left: mean exact match against search count for completed rollouts; all three shapes peak at two searches. Right: the search-count distribution per shape with the per-shape abort rate; EM-only places far more mass in the aborted turn-budget bin, the main source of its lower held-out exact match.
Run
NQ
TriQA
PopQA
HotQA
2Wiki
MuSiQue
Bambgl
avg EM
F1-only s42
0.301
0.492
0.319
0.309
0.287
0.130
0.352
0.313 @110
F1-only s43
0.301
0.489
0.321
0.259
0.274
0.093
0.296
0.291 @210
F1-only s44
0.363
0.531
0.370
0.359
0.285
0.161
0.392
0.352 @310
F1+format s42
0.319
0.500
0.335
0.311
0.313
0.117
0.360
0.322 @100
F1+format s43
0.335
0.533
0.352
0.322
0.296
0.092
0.264
0.313 @210
F1+format s44
0.310
0.521
0.340
0.307
0.278
0.121
0.288
0.309 @270
Appendix
Table 3: Per-benchmark EM at each run’s seed-horizon best checkpoint (step in the avg-EM column). MuSiQue is the single in-distribution set; the other six are out-of-distribution, so the generalisation claim rests on them.
Reward shape
seed 42
seed 43
seed 44
mean
range
F1-only
0.402
0.369
0.407
0.393
0.038
F1+format
0.413
0.370
0.390
0.391
0.043
EM-only
0.368
0.333
0.319
0.340
0.049
Appendix
Table 4: Matched horizon, average token-F1 per reward shape at each run’s best-F1 checkpoint, showing the same ordering on the dense metric.
Reward shape
seed 42
seed 43
seed 44
mean
F1-only
0.394
0.355
0.397
0.382
F1+format
0.378
0.359
0.355
0.364
EM-only
0.334
0.293
0.340
0.323
Appendix
Table 5: Substring-cover accuracy (ACC) at each run’s seed-horizon best-EM checkpoint. EM-only is worst on the most lenient metric as well.