Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward methods can incur substantial judging cost. We introduce RankBuffer, which maintains an ordered, query-specific buffer of previously judged responses as a reusable quality scale. Each rollout is first inserted into an anchor interval through an independent coarse judgment, after which only rollouts assigned to the same interval undergo local fine ranking. The resulting complete order is converted into bounded rank rewards, while boundary expansion, local refinement, and inactive-anchor pruning adapt the buffer as the policy evolves. Across four open-ended benchmarks, RankBuffer consistently outperforms all pointwise baselines. It also achieves nearly on-par performance with the strongest ranking-based reward baseline while substantially reducing judging cost. Ablations demonstrate the importance of both local fine ranking and anchor response content, while buffer analyses show that rollout-derived anchors progressively extend and refine the covered quality scale. These results establish response reuse as an effective approach to efficient relative reward construction.
Figures & tables
Figure 1: Reward construction with pointwise reward methods, conventional ranking-based reward methods, and RankBuffer. Pointwise rewards are assigned independently, so reward values within a rollout group may cluster or tie and obscure relative differences. Ranking-based reward methods first obtain a complete order and map it to evenly spaced reward levels that preserve the relative quality of the rollouts. RankBuffer produces the same structured reward signal using reusable anchors and selective fine ranking.
Figure 2: RankBuffer workflow. An ordered, query-specific buffer enables parallel coarse ranking, followed by local fine ranking within intervals containing multiple rollouts. The resulting complete ranking determines the rewards and GRPO advantages. Boundary expansion, local refinement, and inactive-anchor pruning update the buffer for subsequent iterations.
Table 3Figure 4
Method
AE2
AH-v2
WB-v2
Writing
Avg.
Tokens(M)
Calls(K)
RankBuffer
82.80
30.20
48.79
67.27
57.27
88.8M
36.9K
RankBuffer*
79.69
26.70
46.95
66.71
55.01
68.0M
36.2K
RankBuffer-Bootstrap
79.50
28.40
47.31
66.15
55.34
76.5M
28.4K
RankBuffer-Bootstrap*
81.61
29.60
47.47
67.48
56.54
60.7M
28.1K
Table 2: Effect of buffer initialization and maintenance. An asterisk (*) denotes static buffers.
Method
AE2
AH-v2
WB-v2
Writing
Avg.
Tokens(M)
Calls(K)
RankBuffer
82.80
30.20
48.79
67.27
57.27
88.8M
36.9K
RankBuffer-Bootstrap
79.50
28.40
47.31
66.15
55.34
76.5M
28.4K
LLM
81.74
30.50
47.34
66.83
56.60
96.5M
41.6K
LLM + look-ahead
79.07
28.40
46.31
66.13
54.98
111.0M
42.1K
Table 3: Comparison of RankBuffer, RankBuffer-Bootstrap, and LLM-based buffer management.
Figure 6: Dynamic buffer evolution: (a) anchor counts by source; (b) source composition along the quality scale.
Init.
K0
AE2
AH-v2
WB-v2
Writing
Avg.
Tokens (M)
Calls (K)
Nectar
2
80.99
30.00
46.73
66.69
56.10
79.3M
36.4K
4
82.80
30.20
48.79
67.27
57.27
88.8M
36.9K
7
81.06
29.60
48.02
66.92
56.40
114.1M
36.8K
Bootstrap
2
80.62
27.90
47.55
66.79
55.72
66.5M
28.0K
4
79.50
28.40
47.31
66.15
55.34
76.5M
28.4K
7
79.50
33.50
48.58
67.39
57.24
92.4M
28.7K
Table 4: Effect of the initial anchor count.
Variant
AE2
AH-v2
WB-v2
Writing
Avg.
Tokens (M)
Calls (K)
RankBuffer
82.80
30.20
48.79
67.27
57.27
88.8M
36.9K
Coarse-only
78.30
27.60
45.56
65.89
54.34
74.1M
31.0K
Table 5: Coarse-only ablation.
Initialization
Variant
AlpacaEval 2
Nectar
RankBuffer
82.80
Order-only
51.93
Bootstrap
RankBuffer
79.50
Order-only
54.53
Table 6: Order-only ablation.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Training
Queries per training step
20
Rollouts per query ( G )
8
Policy learning rate
1×10−6
Maximum prompt length
4,096 tokens
Maximum response length
4,096 tokens
Appendix
Table 7: Main implementation settings.
Method
AE2
AH-v2
WB-v2
Writing
Avg.
Tokens
Calls
GRPO
72.73
27.50
46.57
65.42
53.06
25,035,966
31,374
GDPO
64.35
20.10
45.14
64.94
48.63
86,937,684
144,236
Dr. GRPO
56.15
18.90
44.32
63.37
45.60
21,936,802
31,649
DAPO
74.22
26.60
46.83
66.13
53.44
24,078,025
31,231
GPG
53.04
16.90
45.78
64.98
45.18
20,915,602
31,619
Tournament-GRPO
80.75
29.50
48.87
68.16
56.82
157,313,284
84,000
Appendix
Table 8: Complete effectiveness and judge-cost comparison. An asterisk (*) denotes static buffers.
Open-ended query generation lacks standard answers, thus necessitating an effective reward mechanism. Pointwise scoring rubrics provide limited information about the relative quality of sample answers under the same prompt; merging multiple rubric judgments into a single score may also mask the differences between these answers. We propose MatrixReward, which constructs rewards from a rollout-by-rubric win-rate matrix obtained by comparing every pair of sampled responses under each rubric. The spread of each matrix column captures how strongly that rubric distinguishes the current rollouts, while correlations between columns reveal rubric repetition; together, these statistics yield data-dependent rubric weights. We combine these weights with the prior weights of rubrics. After column normalization and weighting, the observed per-rubric maxima and minima define positive and negative ideal profiles. Each rollout's distances to these two ideals determine its relative-closeness quality reward. Evaluated using Qwen3-8B on four open-ended query-answering benchmarks, MatrixReward achieves an average score of 63.02, outperforming the strongest baseline by approximately 2.0%. These results support the idea that matrices derived from relative comparisons can be used to construct rewards more reasonably for open-ended generative reinforcement learning.
Zihan Shen, Qi Liu, Zixuan Yang +4
Zhejiang University · Qwen Business Unit of Alibaba · Renmin University of China
Reinforcement learning in open-ended long-form generation is challenging because reliable reference answers and automatic metrics are often unavailable. Existing rubric-based methods typically rely on pointwise LLM-as-a-judge scoring, but absolute scores are difficult to calibrate across complex responses, may provide weak discrimination among same-query rollouts, and can become saturated during optimization. We propose Tournament-GRPO, a group-wise reward framework that converts rubric-guided LLM judgments into relative rewards through repeated multi-round tournaments among same-query rollouts. Tournament-GRPO compares candidates within groups, accumulates tournament outcomes, and normalizes them into group-wise rewards for GRPO training. Experiments on Deep Research Bench show that Tournament-GRPO consistently outperforms existing reward-design baselines, achieving a 4.52-point overall-score improvement over the strongest baseline. Further analyses show that tournament rewards provide a favorable effectiveness--efficiency trade-off and that tournament design affects training dynamics. These results suggest that rubric-guided tournament comparison provides an effective reward signal for reinforcement learning in open-ended long-form generation.
Zixuan Yang, Yiqun Chen, Wei Yang +7
Renmin University of China · University of Southern California · Zhejiang University +1
Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models. However, despite their strong capabilities in response ranking, generative reward models have not realized their potential in reinforcement learning (RL). Our analysis reveals that this limitation arises from a mismatch between the comparative nature of generative reward modeling and the scalar scoring paradigm adopted by existing RL algorithms. To bridge this gap, we propose a Ranking-based Reward Construction (RRC) approach, which enables generative reward models to provide more effective RL learning signals by deriving rewards from relative preference rankings. RRC introduces two complementary strategies: self-competitive ranking, which exploits comparisons among sampled responses, and anchor-guided ranking, which enables scalable ranking-based reward construction with a small set of reference responses. Experiments across open-ended chat and reasoning benchmarks demonstrate that RRC substantially improves RL training with generative reward models, achieving consistent gains over existing reward construction approaches. Our code can be found at https://github.com/wangclnlp/RRC.
Chenglong Wang, Ziming Zhu, Yifu Huo +9
School of Computer Science and Engineering, Northeastern University, Shenyang, China · 2NiuTrans Research, Shenyang, China · 3Independent Researcher, Beijing, China +1