Reinforcement learning (RL) changes not only what language models say, but also how much they say, often increasing response length at the cost of token efficiency. Controlling this length growth is particularly challenging in open-ended RL because (i) response length is entangled with quality, (ii) open-ended tasks lack a natural success boundary for deciding when efficiency should be prioritized, and (iii) dense, graded rewards often yield small within-group quality margins, making quality-induced advantages especially sensitive to reward-level length shaping, which can perturb their magnitudes and even reverse their signs. We therefore adopt an asymmetric principle: quality should determine the direction of reinforcement, while length should only shape its magnitude. We instantiate this principle with Quality-Gated Length Advantage Shaping (QGLAS), which first computes advantages from quality rewards alone, then adds bounded bonuses only to shorter positive-advantage responses, leaving all other advantages unchanged. The bonus strength is further adapted to within-group quality separation, allowing conciseness to matter more when quality-favored responses are similar and less when their quality differences are clear. Across different model families, open-ended benchmarks, and reward sources, QGLAS consistently achieves a stronger quality--length trade-off than representative baselines. At approximately 30% compression, QGLAS retains 98.4--102.0% of the macro-average quality gains achieved by quality-only RL over the base model, compared with 68.3--75.5% for these baselines at comparable compression.
Figures & tables
Method
IFBench
Hard Prompts
Creative Writing
Macro Average
Relative Metrics
Score ↑
Δ Score ↑
Score ↑
Δ Score ↑
Score ↑
Δ Score ↑
Score ↑
#Tok. ↓
QGR ↑
CR ↑
Qwen3-4B
Base
29.22
–
15.63
–
16.53
–
20.46
3792
0.0%
9.9%
NoBonus
26.89
0.00
21.03
0.00
48.20
0.00
32.04
4207
100.0%
0.0%
GR 3
25.07
-1.82
17.07
-3.96
45.47
-2.73
29.20
2824
75.5%
32.9%
GRLC
24.78
-2.11
20.13
-0.90
42.50
-5.70
29.14
2877
74.9%
31.6%
Table 1: Main results at approximately matched compression. Δ Score denotes the score change relative to quality-only RL. QGR and CR are computed from the unrounded macro averages. Best quality results among length-controlled methods are shown in bold .
Method
IFB. ↑
Hard ↑
Creative ↑
Avg. Score ↑
Avg. #Tok. ↓
QGR ↑
CR ↑
NoBonus
26.89
21.03
48.20
32.04
4207
100.0%
0.0%
QGLAS
25.96
20.40
50.47
32.28
2845
102.0%
32.4%
A. Structural constraints (adaptive scaling retained)
w/o All Structural Constraints
23.33
16.13
42.77
27.41
2895
60.0%
31.2%
w/o Advantage-Level Shaping
25.33
19.93
46.67
30.64
2840
87.9%
32.5%
w/o Positive-Only Gating
24.89
17.83
44.57
29.10
2799
74.6%
33.5%
Table 2: Ablation of QGLAS design choices on Qwen3-4B at approximately matched compression. Each ablation modifies full QGLAS independently. IFB. denotes IFBench, Hard denotes Hard Prompts, and Creative denotes Creative Writing. Best quality results among length-controlled variants are shown in bold.
Training Reward
Method
IFB. ↑
Hard ↑
Creative ↑
Avg. Score ↑
Avg. #Tok. ↓
QGR ↑
CR ↑
Learned RM
NoBonus
26.89
21.03
48.20
32.04
4207
–
–
QGLAS
25.96
20.40
50.47
32.28
2845
102.0%
32.4%
Rubric-based Judge
NoBonus
33.11
19.07
28.87
27.02
4125
–
–
QGLAS
33.44
18.93
28.57
26.98
2960
99.4%
28.2%
LLM-as-a-Judge
NoBonus
32.88
17.13
23.50
24.50
4248
–
–
QGLAS
32.67
16.87
23.90
24.48
3016
99.4%
29.0%
Table 3: Robustness across training quality rewards on Qwen3-4B. IFB. denotes IFBench, Hard denotes the Hard Prompts subset of Arena-Hard-v2, and Creative denotes its Creative Writing subset.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
RL algorithm
GSPO
Optimizer
Adam
Learning rate
1×10−6
Learning-rate schedule
Constant
Adam (β1,β2)
(0.9,0.98)
Weight decay
0.1
Appendix
Table 4: Training configuration for the main experiments.
Reward source
# hi>0
Clipped
Clip rate
Learned RM
111,418
801
0.719%
Rubric-based Judge
118,504
936
0.790%
LLM-as-a-Judge
101,569
1,306
1.286%
Pooled
331,491
3,043
0.918%
Appendix
Table 5: Activation frequency of the clipping threshold c=0.5 in equation 4 . Statistics are computed from stored rollout groups from steps 100–900 of QGLAS (βmin,βmax)=(0.3,0.6) training. The denominator includes only responses with hi>0 ; “Clipped” denotes responses whose pre-clipping relative-shortening term exceeds c .
IFBench
Hard Prompts
Creative Writing
Macro Average
Method
Score ↑
#Tok. ↓
Score ↑
#Tok. ↓
Score ↑
#Tok. ↓
Score ↑
#Tok. ↓
NoBonus
26.89±0.29
2834±65
21.03±0.32
7149±123
48.20±0.79
2638±61
32.04±0.47
4207±83
Run 1 (seed 42)
26.78
2816
20.90
7105
47.90
2620
31.86
4180
Run 2 (seed 43)
27.22
2780
21.40
7054
49.10
2588
32.57
4141
Run 3 (seed 44)
26.67
2906
20.80
7288
47.60
2706
31.69
4300
GR 3
25.07±0.34
1403±48
17.07±0.38
5296±116
45.47±0.83
1774±60
29.20±0.51
2824±74
Appendix
Table 6: Repeated Qwen3-4B results across three training seeds. Mean rows report mean ± sample standard deviation across seeds. Each trained policy is evaluated three times, with evaluation results averaged within each training run.
IFBench
Hard Prompts
Creative Writing
Macro Average
Relative Metrics
Method
Score ↑
#Tok. ↓
Score ↑
#Tok. ↓
Score ↑
#Tok. ↓
Score ↑
#Tok. ↓
QGR ↑
CR ↑
Base
29.22
2548
15.63
6823
16.53
2004
20.46
3792
0.0%
9.9%
NoBonus
26.89
2834
21.03
7149
48.20
2638
32.04
4207
100.0%
0.0%
QGLAS (0.1,0.2)
26.56
2203
20.87
6382
52.63
2301
33.35
3629
111.3%
13.8%
QGLAS (0.2,0.4)
26.11
1892
20.50
5596
53.10
1914
33.24
3134
110.3%
25.5%
QGLAS (0.3,0.6)
25.96
1432
20.40
5379
50.47
1723
32.28
2845
102.0%
32.4%
Appendix
Table 7: Full quality–length trade-off results on Qwen3-4B. Each benchmark reports quality score and average response length. Parenthetical values denote (βmin,βmax) for QGLAS, α for GR 3 , and (λ,β) for GRLC. QGR and CR are computed according to equation 10 . Higher values are preferred for both metrics.
IFBench
Hard Prompts
Creative Writing
Macro Average
Relative Metrics
Variant
Setting
Score ↑
#Tok. ↓
Score ↑
#Tok. ↓
Score ↑
#Tok. ↓
Score ↑
#Tok. ↓
QGR ↑
CR ↑
Base
–
29.22
2548
15.63
6823
16.53
2004
20.46
3792
0.0%
9.9%
NoBonus
–
26.89
2834
21.03
7149
48.20
2638
32.04
4207
100.0%
0.0%
QGLAS
(0.3,0.6)
25.96
1432
20.40
5379
50.47
1723
32.28
2845
102.0%
32.4%
A. Structural constraints (adaptive scaling retained)
w/o All Structural Constraints
Original
23.00
734
15.43
3498
34.70
1032
24.38
1755
33.8%
58.3%
Appendix
Table 8: Full QGLAS ablation results on Qwen3-4B. Structural ablations include both original and approximately compression-matched configurations; parenthetical values in “Matched” rows denote the scalar shaping-strength setting. Adaptive ablations use the fixed statistics shown in the Setting column. Each ablation modifies full QGLAS independently. Best quality results among length-controlled variants are bold.
Pair set
Learned RM
Rubric-based Judge
LLM-as-a-Judge
Overall
3.48%
3.02%
1.91%
Smallest-gap 20%
12.30%
5.41%
5.09%
20–40%
4.02%
5.10%
2.57%
40–60%
0.95%
2.34%
0.69%
60–80%
0.14%
1.43%
0.73%
Largest-gap 20%
0.00%
0.83%
0.47%
Appendix
Table 9: Pairwise ranking-reversal rates among positive-advantage responses. “Overall” is computed over all eligible pairs. Gap-quintile rows report rates conditional on each 20% bin of the pre-shaping quality-advantage gap; smaller-gap bins correspond to more weakly separated quality preferences.
Training Reward
∣P∣=0
∣P∣=1
∣P∣≥4
Mean ∣P∣
Learned RM
0.30%
0.10%
99.52%
8.16
Rubric-based Judge
0.08%
0.12%
99.02%
8.67
LLM-as-a-Judge
0.61%
0.20%
97.54%
8.46
Appendix
Table 10: Statistics of quality-favored responses per rollout group.
Comparison
Seed 42
Seed 43
Seed 44
Mean ± Std.
Hard Prompts
QGLAS vs. GR 3
57.3
59.2
59.7
58.7±1.3
QGLAS vs. GRLC
54.3
56.2
54.9
55.1±1.0
Creative Writing
QGLAS vs. GR 3
64.3
67.2
66.1
65.9±1.5
QGLAS vs. GRLC
71.3
68.8
70.3
70.1±1.3
Appendix
Table 11: Direct Arena-Hard-v2 pairwise evaluation at matched compression. Policies trained with the same seed are compared directly. Scores above 50 indicate an overall preference for QGLAS. The final column reports mean ± sample standard deviation across seeds.
Hard Prompts
Creative Writing
Method
Standard ↑
Length Ctrl. ↑
Standard ↑
Length Ctrl. ↑
Base
15.63
15.03
16.53
18.93
NoBonus
21.03
16.60
48.20
46.20
GR 3
17.07
14.07
45.47
44.47
GRLC
20.13
16.93
42.50
41.17
QGLAS
20.40
17.30
50.47
48.33
Appendix
Table 12: Standard and length-controlled Arena-Hard-v2 evaluation on Qwen3-4B. Length-controlled scores use Arena-Hard-v2’s built-in length feature in the pairwise Bradley–Terry aggregation. Best results among the length-controlled RL methods are shown in bold.
GPT-4.1
Qwen3.5-397B-A17B-FP8
Method
Hard ↑
Creative ↑
Hard ↑
Creative ↑
NoBonus
21.03 (+0.00)
48.20 (+0.00)
14.33 (+0.00)
59.87 (+0.00)
GR 3
17.07 (−3.96)
45.47 (−2.73)
11.13 (−3.20)
56.40 (−3.47)
GRLC
20.13 (−0.90)
42.50 (−5.70)
12.17 (−2.16)
54.20 (−5.67)
QGLAS
20.40 (−0.63)
50.47 (+2.27)
14.10 (−0.23)
60.43 (+0.56)
Appendix
Table 13: Arena-Hard-v2 scores under two LLM judges. Values in parentheses denote differences relative to NoBonus under the same judge. Despite different absolute score scales, both judges produce the same ordering of the four RL methods.
Recent large reasoning models often develop long chain-of-thought responses during reinforcement learning (RL), resulting in high inference latency and deployment cost. Existing methods for response length control typically rely on explicit length penalties or additional control modules, which require careful tuning and may compromise reasoning quality. We propose Quadrant-weighted Sampling for Length-aware Policy Optimization (QLPO), a simple resampling-based variant of GRPO that introduces implicit length control without modifying the reward function. QLPO first over-generates candidate responses and then resamples the training group by preserving the empirical correct/incorrect ratio while favoring short correct responses and long incorrect responses. This reshapes the training distribution and implicitly encourages shorter model outputs. Across models ranging from 1.5B to 32B parameters, including both base models and strong reasoning models, QLPO consistently improves the accuracy-length trade-off. It reduces response length by 30% to 70% while preserving reasoning performance. These results suggest that structured resampling provides an effective and robust approach to efficient reasoning.
Siwei Chen, Siqi Chen, Xupeng Miao +1
School of Computer Science, Peking University · Beijing Key Laboratory of Software and Hardware Cooperative Artificial Intelligence Systems · Department of Electronic Engineering, Tsinghua University +1
Post-training has split large language model (LLM) alignment into two largely disconnected tracks. Online reinforcement learning (RL) with verifiable rewards drives emergent reasoning on math and code but depends on a programmatic verifier that cannot reach open-ended tasks, while preference optimization handles open-ended generation yet forgoes the continuous exploration that powers online RL. Closing this gap requires a verifier for open-ended quality, but a scalar reward model is the wrong shape for the job. Quality is multi-dimensional, and any scalar score is an incomplete proxy that lets online RL collapse onto whichever axis the score is most sensitive to. We turn instead to the General Preference Model (GPM), which embeds responses into k skew-symmetric subspaces and represents preference as a structured, intransitivity-aware comparison. Building on this, we propose General Preference Reinforcement Learning (GPRL), which carries the k-way structure through to the policy update. GPRL computes per-dimension group-relative advantages, normalizes each on its own scale so no axis can dominate, and aggregates them with context-dependent eigenvalues. The same structure powers a closed-loop drift monitor that detects single-axis exploitation and corrects it on the fly by reweighting dimensions and tightening the trust region. Starting from Llama-3-8B-Instruct, GPRL reaches a length-controlled win rate of 56.51% on AlpacaEval~2.0 while also outperforming SimPO and SPPO on Arena-Hard, MT-Bench, and WildBench by resisting reward hacking across extended training runs.
Muhammad Umer, Muhammad Ahmed Mohsin, Ahsan Bilal +5
Recent work trains reasoning models with length penalties to curb overthinking and cut inference cost. We show that these penalties make the chain of thought less monitorable. A length-compressed model still lets misleading hints steer its answers, but it less often verbalizes their influence. We train Qwen3-4B and Qwen3-14B with reinforcement learning under length penalties targeting 60% down to 30% of baseline chain-of-thought length, then evaluate them with nine types of biasing hints on held-out MMLU-Pro-R and four transfer benchmarks. A chain is faithful when an LLM monitor can tell from it that the hint influenced the answer. At the 30% target, accuracy stays near baseline and wrong-answer hints switch answers as often as before. Yet faithfulness drops on every evaluation set for both models, by 39% for Qwen3-14B and 35% for Qwen3-4B on MMLU-Pro-R. A control trained with the same correctness and format rewards but no length penalty leaves faithfulness intact or raises it. Shortening alone does not explain the drop. Compressed chains mention the hint 7 to 35 percentage points less often than the uncompressed model's chains shortened to the same length by random sentence deletion, across both model sizes and all five evaluation sets. Length penalties therefore trade monitorability for inference cost by removing the evidence monitors depend on.