Group-relative policy optimization relies on reward-derived advantages and sequence-level likelihood weights, both of which can be sensitive to localized outliers. Extreme rewards can collapse the contrast among clean responses after group normalization, while token-level log-ratio perturbations can alter sequence weights and clipping decisions. We introduce RoVR-GSPO, a dual-channel robust optimizer that addresses these failure modes separately. Its reward channel combines robust reference estimation with bounded residual credit, while its ratio channel uses differentiable SoftRoVR aggregation to construct robust sequence weights. We provide stability and efficiency analyses for both channels. Experiments on mathematical reasoning, long-context summarization, and tool-call annotation show consistent improvements over GSPO, while controlled perturbation studies demonstrate stronger robustness to reward contamination and token-ratio anomalies.
Figures & tables
Figure 1: Overview of RoVR-GSPO. The reward channel first estimates a robust reference, then applies the bounded residual map and mapped RMS to form Acredit . The ratio channel converts token-level log-ratio sequences into differentiable robust sequence weights via SoftRoVR. These two signals jointly determine the policy update.
Model
Dataset
Rule reward
InternLM2 reward
GSPO
RoVR-GSPO
GSPO
RoVR-GSPO
Qwen3-4B-Base
MATH500
77.00±0.43
77.47±0.57
74.53±0.41
75.80±0.28
Gaokao2023en
39.74±0.21
45.37±0.95
38.53±0.32
42.08±0.37
MinervaMath
19.36±0.46
21.69±1.04
15.81±0.30
16.67±0.34
AIME2025
14.44±1.57
15.56±1.57
8.89±1.57
15.56±1.57
AMC23
62.50±2.04
63.33±1.18
56.67±1.18
64.17±3.12
Table 1: Mathematical-reasoning accuracy (%), mean ± standard deviation over three runs. Bold marks the higher mean for each setting.
Model
Method
ROUGE-1
ROUGE-2
ROUGE-L
Qwen3-4B-Base
GSPO
0.309
0.142
0.166
RoVR-GSPO
0.333
0.164
0.175
OLMo-3-7B-Instruct
GSPO
0.552
0.189
0.221
RoVR-GSPO
0.553
0.195
0.229
Table 2: GovReport ROUGE point estimates on a 0–1 scale; bold marks the higher score.
Method
MATH500
Gaokao
Minerva
AIME25
AMC23
GSPO
74.53±0.50
38.53±0.40
15.81±0.37
8.89±1.92
56.67±1.44
RoVR-Adv
74.73±0.81
40.61±0.40
15.18±0.23
12.22±1.92
56.67±1.44
SoftRoVR
74.53±0.23
40.87±1.28
16.42±0.34
14.44±1.93
59.17±1.44
RoVR-GSPO
75.80±0.35
42.08±0.45
16.67±0.42
15.56±1.93
64.17±3.82
Table 3: Channel-ablation accuracy (%) on mathematical reasoning. Values are reported as mean ± standard deviation over three independent runs; bold marks the highest mean in each column.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: GovReport training dynamics for Qwen3-4B-Base (left pair) and OLMo-3-7B-Instruct (right pair). Each pair shows training reward followed by validation reward. Blue circles denote GSPO, and red squares denote RoVR-GSPO, labeled GSPO (RoVR) in the figure. The green fill visualizes the reward gap.
Figure 3: Training dynamics for the channel ablation: reward, actor entropy, and actor gradient norm. Orange, blue, and red denote RoVR-Adv, SoftRoVR, and RoVR-GSPO; dashed gray denotes GSPO.
Figure 4: Localized token-ratio stress test at recorded step 18 . Curves show means over 192 prompt groups ( 3,072 responses); shaded regions show pointwise 95% prompt-cluster bootstrap intervals. SoftRoVR attenuates log-weight displacement for both isolated spikes and contiguous bursts. For strong spikes, the paired flip-rate difference favors SoftRoVR; for the positive 20% burst, its interval includes zero.
Perturbation
Sign
Metric
GSPO
SoftRoVR
Difference [95% CI]
Spike 16σℓ
+
Dq
1.097
0.808
0.288 [0.256, 0.324]
Spike 16σℓ
+
Flips (%)
22.526
19.362
3.164 [2.246, 4.082]
Burst 20%
+
Dq
49.812
6.258
43.554 [42.112, 44.963]
Burst 20%
+
Flips (%)
47.591
47.956
-0.365 [-0.944, 0.221]
Spike 16σℓ
−
Dq
1.097
0.811
0.286 [0.254, 0.321]
Spike 16σℓ
−
Flips (%)
21.061
17.370
3.691 [2.949, 4.453]
Appendix
Table 4: Paired offline token-ratio stress test at step 18 using 192 prompt groups and 3,072 responses. Dq is reported in 10−3 units; for Flips, differences are reported in %. Intervals are pointwise 95% prompt-cluster bootstrap intervals.
Figure 5: Offline sensitivity to additive reward contamination. One of G=16 rewards is perturbed by ±ασ ; curves summarize 5,000 prompt groups and shaded regions denote bootstrap 95% intervals. The RoVR reference (labeled M-center) limits location displacement, while bounded-credit normalization additionally limits scale inflation and preserves more clean-response advantage contrast.
Estimator
Scale inflation
Clean contrast retention
Advantage RMS deviation
GRPO/GSPO
27.400
0.036
0.982
Robust center
28.194
0.035
0.980
Acredit
1.911
0.520
0.543
Appendix
Table 5: Estimator sensitivity at the strongest offline stress level ( α=16 ). Scale inflation and clean contrast retention are ideal at 1 ; advantage RMS deviation is ideal at 0 . Values are medians across the same 5,000 complete prompt groups used in Figure 5 .
Figure 6: Negative-direction sensitivity analysis at step 18 ( 192 prompt groups, 3,072 responses). Bands are pointwise 95% prompt-cluster bootstrap intervals. The perturbations are −cσℓ for spikes and −8σℓ for bursts. SoftRoVR reduces Dq throughout the tested nonzero settings. Its clipping benefit is larger for strong spikes and short bursts than for the longest burst.
Figure 7: Paired differences across 35 complete steps for spike 16σℓ and burst 20% stress, using 64 prompt groups per step. Lines connect the recorded estimates without smoothing; shaded bands are pointwise 95% prompt-cluster bootstrap intervals. Positive values favor SoftRoVR. Effects on log-weight displacement are consistent across the recorded stages; long-burst clipping effects are stage-dependent.
Estimator
Local statistic
Outer aggregation
Bounded local score
Block protection
Clean outer factor
Arithmetic mean
Mean
None
–
–
1
Global M -center
Bounded-score M -center
None
✓
–
–
MOM
Arithmetic block mean
Median
–
✓
π/2
VRMOM
Arithmetic block mean
Composite quantile
–
✓
VK→π/3
Robust MOM
Robust local M -center
Median
✓
✓
π/2
RoVR
Robust local M -center
Tie-neutral composite quantile
✓
✓
VK→π/3
Appendix
Table 6: Comparison of representative location/reference estimators. RoVR combines a bounded local M -estimator with tie-neutral composite-quantile aggregation while retaining explicit block protection.
Variant
RoVR reference
Numerator
Scale
Role
Aicenter
RoVR(R1:G)
Ri−θR
Raw RMS
Center-only ablation
AiLOO
RoVR(R−i)
Ri−θR,−i
S−i
Leave-one-out theoretical reference
Aicredit
RoVR(R1:G)
χκ(Ri−θR)
sχ
Default robust advantage ; bounded leverage and fidelity control
Appendix
Table 7: RoVR-based advantage constructions used in our analysis and algorithm.
K
VK (theory)
Oracle correction (emp.)
Outer median (emp.)
1
1.5708
1.5812
1.5649
3
1.1680
1.1802
1.5649
5
1.1034
1.1102
1.5649
9
1.0691
1.0747
1.5649
15
1.0564
1.0614
1.5649
31
1.0498
1.0540
1.5649
Appendix
Table 8: Oracle outer-factor validation. The empirical factors use 20,000 trials with B=101 and block size n=32 . The median factor is shown for reference; π/2 and π/3 are the clean asymptotic reference values.
Figure 8: Oracle outer-factor validation. The oracle composite-quantile estimator follows VK closely, while the outer median remains near π/2 .
Figure 9: Sampling variance across clean and contaminated scenarios. Error bars are 95% bootstrap Monte Carlo intervals over 3,000 estimator trials. Under contamination, variance alone is not a sufficient robustness metric because it does not include systematic displacement.
Scenario
Mean
Global M
MOM
VRMOM-style
Robust MOM
RoVR
Gaussian
0.007552
0.008088
0.010271
0.008033
0.010865
0.008596
Student- t3
0.007567
0.004274
0.008424
0.007038
0.005871
0.004660
Point contamination
0.007552
0.009483
0.010271
0.008273
0.012848
0.009828
Block contamination
0.007552
0.009486
0.011677
0.009443
0.012323
0.010190
Appendix
Table 9: Sampling variance of the six reference estimators. Each entry is computed from 3,000 trials; the full CSV also reports bootstrap Monte Carlo intervals.
Scenario
Mean
Global M
MOM
VRMOM-style
Robust MOM
RoVR
Point contamination
1.0059
0.2719
1.0067
1.0064
0.2746
0.2695
Block contamination
1.0059
0.2718
0.1171
0.1129
0.1206
0.1170
Appendix
Table 10: RMSE to the uncontaminated center 0 under the fixed +8 contamination. Both contamination geometries modify 16 of 128 observations.
Figure 10: RMSE to the uncontaminated center under the two matched contamination geometries. Error bars are 95% bootstrap Monte Carlo intervals over 3,000 trials.
Figure 11: RMSE sweep for dispersed point corruption. The y-axis is logarithmic to show both mean-like and robust estimators over the same range.
Figure 12: RMSE sweep for one-block coherent corruption. The favorable behavior of block estimators is conditional on the declared block failure geometry.