Reinforcement Learning with Verifiable Rewards (RLVR) methods such as GRPO rely on successful self-generated trajectories, but finite rollout budgets can produce all-fail groups with no reward-based policy-gradient signal. While additional rollouts improve the chance of success at higher cost, successful trajectories missing from one model's rollouts may already have been discovered by another. Indeed, we observe that heterogeneous models often succeed on complementary prompts, creating opportunities for mutual learning without a designated stronger teacher. To exploit this complementarity, we propose GRAFT (Gated Replacement of Answer-Failed groups with peer Trajectories), an off-policy-aware framework that replaces all-fail groups with informative peer groups. GRAFT transfers both successful and unsuccessful peer responses with peer-computed advantages, while controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping. Across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT consistently improves both models over GRPO with the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points in model-level average performance. Stored peer trajectories preserve most of the gains, improving over GRPO by 1.8 points on average without simultaneous co-training.
Figures & tables
Figure 1: Complementary successes persist throughout training. (Left) Illustrative examples where one model solves a prompt its peer fails entirely, so each could learn from the other’s trajectories. (Right) Fraction of all-fail prompts solved by the peer, with means shown as dotted lines.
Figure 2: Overall Framework of GRAFT. GRAFT selects and balances complementary peer groups, then controls their off-policy influence through compatibility weighting, token-level clipping, and peer-last updates.
Method
MATH500
AIME2024
AIME2025
AMC23
Minerva
Avg.
Δ Avg
Updated policy: SmolLM3-3B-Base
GRPO ( n=8 )
72.08 ± 0.52
8.58 ± 1.13
8.50 ± 0.96
46.62 ± 1.49
27.22 ± 0.63
32.60 ± 0.49
–
GRPO ( n=16 )
72.88 ± 0.40
9.08 ± 0.90
9.83 ± 1.34
47.19 ± 1.01
27.90 ± 0.75
33.38 ± 0.25
↑ 0.78
GRPO ( n=32 )
76.68 ± 0.56
11.42 ± 0.81
12.83 ± 1.23
49.56 ± 1.80
28.08 ± 0.57
35.71 ± 0.18
↑ 3.11
Peer policy: Qwen3-1.7B-Base (Pair 1)
HACPO
69.58 ± 0.53
7.08 ± 1.79
10.83 ± 1.69
42.88 ± 1.61
27.17 ± 0.61
31.51 ± 0.43
↓ 1.09
Table 1: Results of cross-model rollout exchange during co-training. Each updated policy uses the same GRPO baselines for both of its peer policies. Bold : best among the n=8 methods [GRPO ( n=8 ), HACPO, SGT, GRAFT] within each peer block; no bold means GRPO ( n=8 ) is best. † / ‡ : GRAFT exceeds all baselines up to GRPO ( n=16 )/( n=32 ), respectively. Δ : change in average score relative to GRPO ( n=8 ).
Figure 3: (a) Average score of Pair 1 vs. total GPU-hours. Error bars: 95% CIs over five evaluation runs; the dashed line traces GRPO with increasing rollout budget, and DS denotes dynamic sampling. GRAFT exceeds GRPO ( n=32 ) by 1.18 at 0.45 × the cost. (b) Replacing the co-trained partner with its stored trajectories across all six blocks. Top: GPU-hours to the reported checkpoint; bottom: gain over GRPO ( n=8 ). Stored trajectories cut compute by 27–76% while keeping a positive gain in every block (84% of the online gain on average).
Table 2: Ablations and alternative designs on Pair 1 (S: SmolLM3-3B, Q: Qwen3-1.7B; Δ : change vs. full GRAFT). (a) removes or randomizes one component at a time; (b) replaces the transfer rule, or fixes only one of which and how while borrowing the other. Variant definitions and per-benchmark scores: Appendices G and F .
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Training epochs
3
Prompt batch size
128
Rollouts per prompt per model
8
Optimization minibatch size
32 prompts
Optimization epochs per rollout batch
1
Maximum prompt / response length
2,048 / 4,096 tokens
Appendix
Table 3: Training hyperparameters. These settings apply to GRAFT and independent GRPO unless otherwise specified. The larger-budget GRPO baselines change only the rollout count.
Figure 4: Validation trajectories used for checkpoint selection. Five-benchmark average vs. training step for each updated model, evaluated under the identical protocol for all methods. Stars mark the selected checkpoints in Table 1 .
Method
MATH500
AIME2024
AIME2025
AMC23
Minerva
Avg.
Δ Avg
Pair 1 SmolLM3-3B-Base ↔ Qwen3-1.7B-Base
SmolLM3-3B-Base
GRPO ( n=8 )
72.21 ± 2.78
7.91 ± 0.58
8.72 ± 2.01
45.08 ± 2.46
26.84 ± 0.45
32.15 ± 1.48
–
GRAFT
75.37 ± 1.30
13.28 ± 1.32
13.22 ± 1.29
50.37 ± 0.98
28.34 ± 1.36
36.12 ± 0.85
↑ 3.97
Qwen3-1.7B-Base
GRPO ( n=8 )
70.69 ± 0.30
9.66 ± 0.43
5.86 ± 0.97
42.60 ± 0.78
28.10 ± 0.35
31.38 ± 0.16
–
Appendix
Table 4: Training stability across three independent runs. We repeat GRPO ( n=8 ) and GRAFT three times for all three model pairs and report the mean and sample standard deviation across runs. Δ denotes the difference in the mean average score relative to GRPO ( n=8 ) within each model block. Bold indicates the better mean performance between GRPO ( n=8 ) and GRAFT.
Figure 5: Per-model score against the GPU-hours charged to that model under the most conservative rule: GRPO pays for its own model only; each model of a cross-model method is charged the full joint-run cost (both models’ training) up to that model’s own selected checkpoint. Error bars are ±1 s.d. over inference seeds.
SmolLM3-3B-Base
Qwen3-1.7B-Base
Both models
Avg.
Best ckpt
3 epochs
Avg.
Best ckpt
3 epochs
Best ckpt
3 epochs
(%)
(GPU-h)
(GPU-h)
(%)
(GPU-h)
(GPU-h)
(GPU-h)
(GPU-h)
GRPO ( n=8 )
32.60 ± 0.49
17.2
17.4
31.20 ± 0.27
11.4
12.5
28.6
29.9
GRPO ( n=16 )
33.38 ± 0.25
21.0
34.8
32.15 ± 0.73
22.1
22.5
43.1
57.3
GRPO ( n=32 )
35.71 ± 0.18
48.7
99.5
32.43 ± 0.65
43.3
45.6
92.0
145.1
GRPO ( n=8 ) + Dynamic Sampling
34.99 ± 0.35
29.4
40.5
32.59 ± 0.33
21.4
27.4
50.8
67.9
Appendix
Table 5: Score and training cost on Pair 1. Avg. is the five-benchmark average at the selected checkpoint ( Table 1 ). GPU-hour columns report the cost up to the selected checkpoint and for the complete three-epoch run. For online cross-model methods, the “Both models / Best ckpt” entry measures the joint run through the later of the two selected checkpoints; for independent GRPO, it sums the two separate checkpoint costs. Stored-trajectory costs exclude the prior runs used to collect peer logs.
Method
MATH500
AIME2024
AIME2025
AMC23
Minerva
Avg.
Δ Avg
Pair 1 SmolLM3-3B-Base ↔ Qwen3-1.7B-Base
SmolLM3-3B-Base
GRPO ( n=8 )
72.08 ± 0.52
8.58 ± 1.13
8.50 ± 0.96
46.62 ± 1.49
27.22 ± 0.63
32.60 ± 0.49
–
GRAFT
76.86 ± 0.32
14.42 ± 1.49
14.50 ± 0.99
50.87 ± 1.60
28.68 ± 0.57
37.06 ± 0.67
↑ 4.46
GRAFT w/ stored traj.
74.53 ± 0.41
12.00 ± 1.62
12.50 ± 0.29
49.94 ± 2.36
27.44 ± 0.69
35.28 ± 0.33
↑ 2.68
Qwen3-1.7B-Base
Appendix
Table 6: Effect of learning from peer-model training trajectories. We compare standard GRPO ( n=8 ), online cross-model rollout exchange (GRAFT), and GRAFT with peer replay, where each model learns from peer trajectories collected from a previous training run rather than from a simultaneously co-trained peer. All methods use n=8 self-rollouts per model. Δ reports the change in average score relative to GRPO ( n=8 ) within each model block.
Variant
MATH500
AIME2024
AIME2025
AMC23
Minerva
Avg.
Δ Avg
SmolLM3-3B-Base
GRAFT (full, δ=0.8 )
76.86 ± 0.32
14.42 ± 1.49
14.50 ± 0.99
50.87 ± 1.60
28.68 ± 0.57
37.06 ± 0.67
–
Removing one component
w/o compatibility gate
68.66 ± 0.40
5.42 ± 1.47
4.92 ± 1.12
39.19 ± 1.66
25.33 ± 0.57
28.70 ± 0.78
↓ 8.36
w/o balanced exchange
75.05 ± 0.25
10.08 ± 1.68
12.50 ± 0.78
49.44 ± 2.20
28.24 ± 0.48
35.06 ± 0.27
↓ 2.00
w/o token-level ratio
75.83 ± 0.58
10.50 ± 1.04
12.50 ± 1.56
48.94 ± 1.80
28.26 ± 0.49
35.21 ± 0.65
↓ 1.85
Appendix
Table 7: Full ablation of GRAFT on Pair 1. Per-benchmark scores behind Table 2 . Each variant changes a single component of the full method while keeping the compute budget fixed at n=8 rollouts per model. Count-matched random variants match the prompt-selection or response-admission count obtained by applying GRAFT to the current rollout batch, separately in each direction, but select uniformly at random. Random admission replaces the compatibility floor while retaining weights min{s,1} . Δ reports the change in average score relative to full GRAFT within each model block.
Method
MATH500
AIME2024
AIME2025
AMC23
Minerva
Avg.
Δ Avg
SmolLM3-3B-Base
GRAFT
76.86 ± 0.32
14.42 ± 1.49
14.50 ± 0.99
50.87 ± 1.60
28.68 ± 0.57
37.06 ± 0.67
–
HACPO
69.58 ± 0.53
7.08 ± 1.79
10.83 ± 1.69
42.88 ± 1.61
27.17 ± 0.61
31.51 ± 0.43
↓ 5.55
HACPO + GRAFT which
67.73 ± 0.43
6.42 ± 0.37
6.42 ± 1.16
40.06 ± 1.64
24.23 ± 0.58
28.97 ± 0.52
↓ 8.09
SGT
75.53 ± 0.51
10.67 ± 1.52
11.50 ± 0.37
49.06 ± 1.51
27.63 ± 0.68
34.88 ± 0.21
↓ 2.18
SGT + GRAFT how
75.05 ± 0.25
10.08 ± 1.68
12.50 ± 0.78
49.44 ± 2.20
28.24 ± 0.48
35.06 ± 0.27
↓ 2.00
Appendix
Table 8: Full results for prior methods with one side replaced, on Pair 1. Per-benchmark results for combinations of prior methods with GRAFT’s prompt selection or peer update. HACPO and SGT rows are from Table 1 ; “SGT + GRAFT how” is identical to “w/o balanced exchange” in Table 7 . Δ reports the change in average score relative to full GRAFT within each model block.
Peer minibatch position
Clipped-token fraction (%)
Update
First
Uniform
Last (ours)
Self-generated
SmolLM3-3B
0.046
0.043
0.045
Qwen3-1.7B
0.068
0.068
0.072
Grafted (peer)
SmolLM3-3B ← Qwen3-1.7B
0.004
0.130
0.277
Qwen3-1.7B ← SmolLM3-3B
0.002
0.248
0.440
Steps with zero clipping on peer tokens (%)
97.6
13.0
0.0
Appendix
Table 9: Token-level clipping under different peer-minibatch orderings on Pair 1, measured over training steps 1–48.
Figure 6: Root shift in Qwen3-1.7B (MATH500). Target GRPO expands g(x)=f(x+5) and uses the expanded polynomial to compute the root sum with Vieta’s formula. The peer instead shifts each root of f by −5 , so their sum decreases by 15 . GRAFT uses the same root shift as the peer, computes the original root sum with Vieta’s formula, and obtains 49−15=34 .
Figure 7: Enforcing the occupancy constraint in Qwen3-1.7B (MATH500). Target GRPO counts all 36=729 lane assignments without enforcing that every lane is occupied. The peer treats an assignment as an onto mapping from the six cars to the three lanes and applies inclusion–exclusion. GRAFT applies the same inclusion–exclusion correction through empty-lane events: it subtracts 3×26 assignments, adds back the three assignments in which only one lane is occupied, and obtains 729−192+3=540 .
Figure 8: Separating the rational and logarithmic factors in Qwen3-1.7B (AIME2025). Target GRPO separates the rational and logarithmic factors, then duplicates the logarithmic ratio when factoring the rational term. The peer separates the rational product from a single logarithmic product that telescopes to 3 . GRAFT uses the same separation with base- 5 logarithms, evaluates the two rational products as 31 and 1/13 , and obtains 31×(1/13)×3=93/13 , so m+n=106 .
Figure 9: Cartesian optimization in SmolLM3-3B (AIME2024). Target GRPO uses the polar parameterization to reduce the objective to 324cosθ−432sinθ , whose maximum is 540 . The peer writes z=x+iy , rewrites the objective as 81x−108y under x2+y2=16 , and solves the constrained problem with Lagrange multipliers. GRAFT starts with the polar parameterization, then switches to z=x+iy , derives the same constrained objective as the peer, and obtains x=12/5 , y=−16/5 , and the maximum 540 from the Lagrange equations.
Figure 10: Recovering the height in SmolLM3-3B (AIME2025). Target GRPO places G on the line containing A,…,F , recognizes the resulting collinearity as an error, but does not recover a nonzero height. The peer keeps a vertical coordinate for G in the distance constraints CG=40 and DG=30 , obtains a height of 24 , and computes the area as 468 with the shoelace formula. GRAFT places G off the line, uses the two distance constraints to recover a height of 24 , and computes the area from BE=39 as 239×24=468 .
Figure 11: Recovering the norm from the squared ratio in SmolLM3-3B (MATH500). Target GRPO maximizes the squared norm ratio f(t) with t=y/x , obtains a maximum of 16 , and sets C=16 without taking the square root. The peer computes the largest eigenvalue of A⊤A as 16 and takes its square root to obtain the operator norm C=4 . GRAFT keeps the scalar optimization used by target GRPO and also takes the square root of the resulting maximum to obtain C=16=4 .
Reinforcement Learning with Verifiable Rewards (RLVR) for language-model reasoning can fail at both extremes of task difficulty: easy prompts often produce all-correct, low-diversity rollout groups with little gradient signal, while hard prompts can produce all-incorrect groups with no positive reward. We introduce ExTra (Exploratory Trajectory Optimization), a GRPO-compatible framework that extracts exploration signals from the model's own rollouts. ExTra combines two mechanisms: (i) a novelty reward that adds embedding-based diversity bonuses after GRPO normalization, rewarding diverse correct solutions; and (ii) entropy-guided prefix regeneration, which scores partial trajectories using entropy signals and continues exploration from promising intermediate steps. Across six mathematical reasoning benchmarks, ExTra improves Qwen3-1.7B over GRPO by about +5 points on pass@1 and +7 points on pass@16, showing that trajectory-level exploration signals can improve both single-sample accuracy and inference-time coverage.
Reinforcement learning with verifiable rewards (RLVR) has emerged as a central paradigm for improving the reasoning capabilities of large language models. Group-based policy optimization methods, such as GRPO, typically allocate a fixed number of rollouts to every prompt. This uniform allocation can be inefficient: it over-allocates compute to prompts whose sampled groups are already saturated while under-exploring prompts for which additional samples may reveal useful correct trajectories. To address this limitation, we introduce hit utility, the posterior probability that at least one rollout in a proposed additional allocation for a prompt will be correct. Building on this notion, we propose Hit-Utility Optimal Rollout Allocation (HORA), a learning-free rollout allocation policy that maximizes total posterior hit utility within each allocation batch. HORA adaptively reallocates rollout budgets while leaving the downstream reward evaluation and group-based advantage estimator unchanged. Across four mathematical reasoning benchmarks and three model scales, HORA preserves comparable Pass@1 and improves Pass@K over compute-matched GRPO in ten of twelve model--benchmark configurations, with one tie and one saturated exception. It is also drop-in compatible with other group-based estimators such as RLOO. Ablation studies indicate that the uniform prior used by HORA is competitive with five prompt-conditioned learned-prior alternatives.
Tao Wang, Shuo Li, Yan Sun +2
University of Pennsylvania · New Jersey Institute of Technology · University of Tennessee, Knoxville
Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting scalability. Voting-based label-free RLVR replace gold supervision with answer-level consensus from model samples. However, collapse arises when the same answer-level signal is used both to estimate rewards and to drive token-level policy optimization, encouraging the model to directly reinforce answer tokens rather than improve reasoning. We propose OM-GRPO, a label-free RLVR framework that decouples reward estimation from policy optimization. OM-GRPO masks gradients on the answer span while retaining answer-level rewards through a soft consensus signal, shifting optimization pressure away from answer tokens. We further introduce Contrast-Augmented Reward, which refines reward estimation via low-cost pairwise comparisons over existing trajectories without additional rollouts. Across diverse reasoning benchmarks and three LLM backbones, OM-GRPO consistently outperforms existing label-free RLVR methods and matches supervised GT-reward training with stable optimization. This stability is particularly beneficial in the Test-Time Training setting, where OM-GRPO surpasses majority voting by 4.24 points.
Yongshi Ye, Liang Zhang, Yidong Chen +2
Xiamen University · Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism