Reinforcement learning from verifiable rewards (RLVR) frequently reuses rollouts across multiple policy updates, increasing the mismatch between the current policy and the data-generating policy. We identify a sign-dependent gradient starvation problem in clipped policy optimization: clipping suppresses under-generated positive responses at the low-importance-weight tail while permitting severely over-generated negative responses to dominate the high-weight tail. To address this, we propose ReSPO (Reshaped Sequence Policy Optimization), which replaces clipping with a smooth, two-branch sequence-level kernel derived from an α-divergence variational objective and an exponential variance-control tilt. The positive branch preserves a nonzero gradient weight for under-generated positive responses, while the negative branch suppresses heavily over-generated negative responses. We demonstrate that ReSPO effectively learns from long positive reasoning trajectories during early training, even when accumulated policy drift relegates them to the low-importance-weight tail. On dense and MoE Qwen3 models, ReSPO accelerates early optimization, improves final training scores, and achieves higher held-out benchmark performance under a rollout reuse, validating our approach on importance-weight tail control in off-policy learning.
Figures & tables
Figure 1 : Kernel shapes and tail limits ( ε=0.2 ). The adjacent table gives the limits as W→(0,∞) . ReSPO preserves positive-branch signal near zero while suppressing both negative tails.
Figure 2 : Normalized training-score curves and summaries. Tables report the peak through step 256 and the last- 128 mean. Early peaks have no error bars; late-mean subscripts are temporal standard deviations. Bold marks each row’s maximum; for late means, it also marks methods whose central value and the maximum each fall within the other’s range.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Qwen3-1.7B
Qwen3-30B-A3B
Model & infrastructure
Training backend
FSDP
Megatron
Tensor parallelism (train)
–
2
Expert parallelism (train)
–
8
Tensor parallelism (rollout)
2
4
Data
Appendix
Table 2 : Training hyperparameters across experimental settings.
Method
Hyperparameter
Value
GRPO
Clip ratio ε
0.2
GSPO
Clip ratio low εlow
3×10−4
Clip ratio high εhigh
4×10−4
VESPO
β(+)
2
λ(+)
3
β(−)
3
Appendix
Table 3 : Method-specific hyperparameters. All methods share the training configuration in Tab. 2 .
N
Entry
Q1
Q2
Q3
Q4
Q5
Q6
Q7
Q8
Q9
Q10
8
Upper bound
355
480
611
761
954
1,263
2,009
4,404
16,383
16,384
GRPO count
1,043
908
820
855
816
809
573
163
102
1,295
GSPO count
1,253
1,087
1,131
1,074
917
775
420
79
46
602
VESPO count
386
466
456
475
539
671
1,037
1,504
1,376
474
ReSPO count
207
429
482
482
605
630
856
1,140
1,360
1,193
16
Upper bound
366
481
604
736
907
1,168
1,785
4,252
16,383
16,384
Appendix
Table 4 : Shared response-length-bin upper bounds (tokens) and pooled response counts for Qwen3-1.7B. Counts aggregate the six evaluation benchmarks; each method contributes 7,384 responses for each N . Q10 contains responses at the 16,384 -token cap.
Figure 4 : Qwen3-1.7B accuracy by shared response-length bin. Q1–Q9 partition uncapped responses using boundaries shared across methods, while Q10 contains responses at the 16,384 -token cap. Columns correspond to N∈{8,16,32} ; gaps in the benchmark panels indicate empty method–benchmark bins.
Statistic
ReSPO
VESPO
Positive responses with logW<−1
22.8%
16.7%
Coefficient mass on logW<−1
38.5%
15.5%
Amplification on logW<−1
1.69×
0.93×
Positive responses in −5<logW≤−2
7.4%
3.7%
Coefficient mass in −5<logW≤−2
14.8%
1.1%
Amplification in −5<logW≤−2
2.00×
0.30×
Appendix
Table 5 : Positive-branch deep-tail allocation on Qwen3-1.7B-Base at N=32 . Values are means over rollout batches 9 – 16 ; amplification is the coefficient-mass share divided by the sequence share.
Mean logW
Mean response length
Batch
ReSPO
VESPO
ReSPO
VESPO
1
−0.021
+0.134
752
829
8
−0.188
−0.064
829
843
16
−0.385
−0.064
1076
915
Appendix
Table 6 : Mean logW and response length on the positive branch during the N=32 diagnostic runs.
ReSPO
VESPO
logW bin
W range
Length
n
Length
n
(−∞,−20]
(0,2.1×10−9]
2155(1.90×)
9
–
0
(−20,−10]
2.1×10−9 – 4.5×10−5
2208(1.95×)
52
–
0
(−10,−5]
4.5×10−5 – 0.007
2227(1.97×)
240
5593(6.35×)
8
(−5,−2]
0.007 – 0.14
1501(1.32×)
1687
1085(1.23×)
592
(−2,−1]
0.14 – 0.37
1199(1.06×)
2753
957(1.09×)
2071
Appendix
Table 7 : Count-weighted positive-response length by sequence-weight bin over rollout batches 1 – 16 of the matched Qwen3-1.7B-Base follow-up runs at N=32 and a 16,384 -token response limit. Parentheses give length relative to the corresponding method’s positive-branch mean over the same window; n is the positive-response count and “–” denotes an empty bin.
Setting
Early peak
Late score
Final score
ReSPO
0.445
0.563±0.021
0.555
ReSPO + R3
0.446
0.582±0.008
0.584
Appendix
Table 8 : Matched Qwen3-30B-A3B ReSPO training diagnostics at N=16 , with and without R3 Routing Replay. All columns except the early peak and final score are time-weighted means over policy steps [896,1024] .
Figure 5 : Detailed training statistics for Qwen3-1.7B-Base on DAPO-MATH. Rows correspond to N∈{8,16,32} . Training accuracy is (critic/score/mean+1)/2 , and the evaluation panel reports AIME25 best@4 accuracy.
Figure 6 : Detailed training statistics for Qwen3-30B-A3B-Base on DAPO-MATH. The training panel reports penalized training accuracy, and the evaluation panel reports AIME25 best@2 accuracy, the largest common best-of- k metric available for all four methods; the other panels follow Fig. 5 .