Reinforcement learning from verifiable rewards (RLVR) frequently reuses rollouts across multiple policy updates, increasing the mismatch between the current policy and the data-generating policy. We identify a sign-dependent gradient starvation problem in clipped policy optimization: clipping suppresses under-generated positive responses at the low-importance-weight tail while permitting severely over-generated negative responses to dominate the high-weight tail. To address this, we propose ReSPO (Reshaped Sequence Policy Optimization), which replaces clipping with a smooth, two-branch sequence-level kernel derived from an α-divergence variational objective and an exponential variance-control tilt. The positive branch preserves a nonzero gradient weight for under-generated positive responses, while the negative branch suppresses heavily over-generated negative responses. We demonstrate that ReSPO effectively learns from long positive reasoning trajectories during early training, even when accumulated policy drift relegates them to the low-importance-weight tail. On dense and MoE Qwen3 models, ReSPO accelerates early optimization, improves final training scores, and achieves higher held-out benchmark performance under a rollout reuse, validating our approach on importance-weight tail control in off-policy learning.
Figures & tables
Figure 1 : Kernel shapes and tail limits ( ε=0.2 ). The adjacent table gives the limits as W→(0,∞) . ReSPO preserves positive-branch signal near zero while suppressing both negative tails.
Figure 2 : Normalized training-score curves and summaries. Tables report the peak through step 256 and the last- 128 mean. Early peaks have no error bars; late-mean subscripts are temporal standard deviations. Bold marks each row’s maximum; for late means, it also marks methods whose central value and the maximum each fall within the other’s range.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Qwen3-1.7B
Qwen3-30B-A3B
Model & infrastructure
Training backend
FSDP
Megatron
Tensor parallelism (train)
–
2
Expert parallelism (train)
–
8
Tensor parallelism (rollout)
2
4
Data
Appendix
Table 2 : Training hyperparameters across experimental settings.
Method
Hyperparameter
Value
GRPO
Clip ratio ε
0.2
GSPO
Clip ratio low εlow
3×10−4
Clip ratio high εhigh
4×10−4
VESPO
β(+)
2
λ(+)
3
β(−)
3
Appendix
Table 3 : Method-specific hyperparameters. All methods share the training configuration in Tab. 2 .
N
Entry
Q1
Q2
Q3
Q4
Q5
Q6
Q7
Q8
Q9
Q10
8
Upper bound
355
480
611
761
954
1,263
2,009
4,404
16,383
16,384
GRPO count
1,043
908
820
855
816
809
573
163
102
1,295
GSPO count
1,253
1,087
1,131
1,074
917
775
420
79
46
602
VESPO count
386
466
456
475
539
671
1,037
1,504
1,376
474
ReSPO count
207
429
482
482
605
630
856
1,140
1,360
1,193
16
Upper bound
366
481
604
736
907
1,168
1,785
4,252
16,383
16,384
Appendix
Table 4 : Shared response-length-bin upper bounds (tokens) and pooled response counts for Qwen3-1.7B. Counts aggregate the six evaluation benchmarks; each method contributes 7,384 responses for each N . Q10 contains responses at the 16,384 -token cap.
Figure 4 : Qwen3-1.7B accuracy by shared response-length bin. Q1–Q9 partition uncapped responses using boundaries shared across methods, while Q10 contains responses at the 16,384 -token cap. Columns correspond to N∈{8,16,32} ; gaps in the benchmark panels indicate empty method–benchmark bins.
Statistic
ReSPO
VESPO
Positive responses with logW<−1
22.8%
16.7%
Coefficient mass on logW<−1
38.5%
15.5%
Amplification on logW<−1
1.69×
0.93×
Positive responses in −5<logW≤−2
7.4%
3.7%
Coefficient mass in −5<logW≤−2
14.8%
1.1%
Amplification in −5<logW≤−2
2.00×
0.30×
Appendix
Table 5 : Positive-branch deep-tail allocation on Qwen3-1.7B-Base at N=32 . Values are means over rollout batches 9 – 16 ; amplification is the coefficient-mass share divided by the sequence share.
Mean logW
Mean response length
Batch
ReSPO
VESPO
ReSPO
VESPO
1
−0.021
+0.134
752
829
8
−0.188
−0.064
829
843
16
−0.385
−0.064
1076
915
Appendix
Table 6 : Mean logW and response length on the positive branch during the N=32 diagnostic runs.
ReSPO
VESPO
logW bin
W range
Length
n
Length
n
(−∞,−20]
(0,2.1×10−9]
2155(1.90×)
9
–
0
(−20,−10]
2.1×10−9 – 4.5×10−5
2208(1.95×)
52
–
0
(−10,−5]
4.5×10−5 – 0.007
2227(1.97×)
240
5593(6.35×)
8
(−5,−2]
0.007 – 0.14
1501(1.32×)
1687
1085(1.23×)
592
(−2,−1]
0.14 – 0.37
1199(1.06×)
2753
957(1.09×)
2071
Appendix
Table 7 : Count-weighted positive-response length by sequence-weight bin over rollout batches 1 – 16 of the matched Qwen3-1.7B-Base follow-up runs at N=32 and a 16,384 -token response limit. Parentheses give length relative to the corresponding method’s positive-branch mean over the same window; n is the positive-response count and “–” denotes an empty bin.
Setting
Early peak
Late score
Final score
ReSPO
0.445
0.563±0.021
0.555
ReSPO + R3
0.446
0.582±0.008
0.584
Appendix
Table 8 : Matched Qwen3-30B-A3B ReSPO training diagnostics at N=16 , with and without R3 Routing Replay. All columns except the early peak and final score are time-weighted means over policy steps [896,1024] .
Figure 5 : Detailed training statistics for Qwen3-1.7B-Base on DAPO-MATH. Rows correspond to N∈{8,16,32} . Training accuracy is (critic/score/mean+1)/2 , and the evaluation panel reports AIME25 best@4 accuracy.
Figure 6 : Detailed training statistics for Qwen3-30B-A3B-Base on DAPO-MATH. The training panel reports penalized training accuracy, and the evaluation panel reports AIME25 best@2 accuracy, the largest common best-of- k metric available for all four methods; the other panels follow Fig. 5 .
Group Relative Policy Optimization (GRPO) is one of the most widely adopted RLVR algorithms for post-training large language models on reasoning tasks. We first show that GRPO admits an equivalent discriminative reformulation, in which policy optimization maximizes the expected score gap between verified positive and negative rollouts. This reformulation reveals two objective-level limitations: likelihood-misaligned surrogate scores, in which clipped ratio-based scores are optimized rather than the sequence likelihoods that govern generation, and score-insensitive credit assignment, in which rollout-level credit does not reflect the current score gaps between positive and negative rollouts. To address these limitations, we propose ConSPO, a Contrastive Sequence-level Policy Optimization method that uses length-normalized sequence log-probabilities as rollout scores and contrasts verified positive rollouts against negative distractors within the same group. ConSPO optimizes a group-wise InfoNCE-style objective to adaptively strengthen updates for poorly separated positives and high-scoring negatives, together with a curriculum-scheduled margin that preserves separation pressure as training progresses. Experiments across diverse settings show that ConSPO outperforms strong baselines on challenging reasoning benchmarks. Code will be released upon paper acceptance.
Feng Zhang, Xinhong Ma, Ziqiang Dong +5
Beijing Institute of Technology · Qwen Business Unit of Alibaba · The Chinese University of Hong Kong, Shenzhen +1
We propose a new perspective on policy optimization: rather than reweighting all samples by their importance ratios, an optimizer should select which samples are trustworthy enough to drive a policy update. Building on this view, we introduce Rejection-Gated Policy Optimization (RGPO), which replaces the importance sampling ratio r_theta = pi_theta / pi_old with a smooth, differentiable acceptance gate alpha_theta(s, a) = g(r_theta(s, a)) in the range [0, 1]. Unlike prior work that applies rejection sampling as a data-level heuristic before training, RGPO elevates rejection to an optimization principle: the gate participates directly in gradient computation and is implicitly updated alongside the policy. RGPO provides a unified framework: the policy gradients of TRPO, PPO, and REINFORCE all correspond to specific choices of the effective gradient weight w(r) = g'(r) * r. We prove that RGPO guarantees finite, bounded gradient variance even when importance sampling ratios are heavy-tailed (where IS variance diverges). We further show that RGPO incurs only a bounded, controllable bias and provides an approximate monotonic policy improvement guarantee analogous to TRPO. RGPO matches PPO in computational cost, requires no second-order optimization, and extends naturally to RLHF-style preference alignment. In online preference fine-tuning of Qwen2.5-1.5B-Instruct on Anthropic HH-RLHF (n = 3 seeds), RGPO uses a dual-ratio gate that anchors learning to both the previous policy and the reference model, achieving a Pareto-dominant outcome: the highest reward among online RL methods (+14.8% vs. PPO-RLHF) and the lowest KL divergence to the reference model (-16.0% vs. PPO-RLHF, -53.1% vs. GRPO).
Reinforcement learning with verifiable rewards (RLVR) is a core post-training recipe for reasoning models, yet pure on-policy learning can be inefficient when useful trajectories are difficult to discover or exploration narrows. Existing self-guided approaches largely reuse capability already available to the current or earlier learner. We instead ask whether learning can also make use of capabilities that emerge later in training: can a model learn from its own future self? We introduce temporal self-distillation, in which a policy receives guidance from a stronger later checkpoint of itself. We hypothesize that the most useful temporal teacher need not be the strongest one: a teacher must provide sufficiently new capability while remaining compatible enough for that capability to be readily transferred, motivating a near-future regime. We study this principle through two complementary mechanisms. Near-Future Policy Optimization (NPO) performs off-policy behavioral transfer using verified future-self trajectories, while Near-Future Policy Distillation (NPD) performs on-policy token-level transfer on learner-generated trajectories. We further introduce AutoNPO, which adaptively determines when temporal guidance is useful and how far to roll back, turning future-self guidance into a repeated self-bootstrap process. Across eight image-text benchmarks, NPO improves GRPO from 60.25 to 62.84 and AutoNPO reaches 63.15, with consistent gains on text-only and video reasoning. Under NPD, a near-future teacher reaches 63.23 after continued RL versus 61.92 with a far-future teacher, despite lower immediate post-distillation performance. Together, these results suggest that effective temporal self-distillation depends not simply on teacher strength, but on a balance between newly acquired capability and learner compatibility.
Chuanyu Qin, Chenxu Yang, Qingyi Si +6
Institute of Information Engineering, CAS · School of Cyber Security, UCAS