Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward methods can incur substantial judging cost. We introduce RankBuffer, which maintains an ordered, query-specific buffer of previously judged responses as a reusable quality scale. Each rollout is first inserted into an anchor interval through an independent coarse judgment, after which only rollouts assigned to the same interval undergo local fine ranking. The resulting complete order is converted into bounded rank rewards, while boundary expansion, local refinement, and inactive-anchor pruning adapt the buffer as the policy evolves. Across four open-ended benchmarks, RankBuffer consistently outperforms all pointwise baselines. It also achieves nearly on-par performance with the strongest ranking-based reward baseline while substantially reducing judging cost. Ablations demonstrate the importance of both local fine ranking and anchor response content, while buffer analyses show that rollout-derived anchors progressively extend and refine the covered quality scale. These results establish response reuse as an effective approach to efficient relative reward construction.
Figures & tables
Figure 1: Reward construction with pointwise reward methods, conventional ranking-based reward methods, and RankBuffer. Pointwise rewards are assigned independently, so reward values within a rollout group may cluster or tie and obscure relative differences. Ranking-based reward methods first obtain a complete order and map it to evenly spaced reward levels that preserve the relative quality of the rollouts. RankBuffer produces the same structured reward signal using reusable anchors and selective fine ranking.
Figure 2: RankBuffer workflow. An ordered, query-specific buffer enables parallel coarse ranking, followed by local fine ranking within intervals containing multiple rollouts. The resulting complete ranking determines the rewards and GRPO advantages. Boundary expansion, local refinement, and inactive-anchor pruning update the buffer for subsequent iterations.
Table 3Figure 4
Method
AE2
AH-v2
WB-v2
Writing
Avg.
Tokens(M)
Calls(K)
RankBuffer
82.80
30.20
48.79
67.27
57.27
88.8M
36.9K
RankBuffer*
79.69
26.70
46.95
66.71
55.01
68.0M
36.2K
RankBuffer-Bootstrap
79.50
28.40
47.31
66.15
55.34
76.5M
28.4K
RankBuffer-Bootstrap*
81.61
29.60
47.47
67.48
56.54
60.7M
28.1K
Table 2: Effect of buffer initialization and maintenance. An asterisk (*) denotes static buffers.
Method
AE2
AH-v2
WB-v2
Writing
Avg.
Tokens(M)
Calls(K)
RankBuffer
82.80
30.20
48.79
67.27
57.27
88.8M
36.9K
RankBuffer-Bootstrap
79.50
28.40
47.31
66.15
55.34
76.5M
28.4K
LLM
81.74
30.50
47.34
66.83
56.60
96.5M
41.6K
LLM + look-ahead
79.07
28.40
46.31
66.13
54.98
111.0M
42.1K
Table 3: Comparison of RankBuffer, RankBuffer-Bootstrap, and LLM-based buffer management.
Figure 6: Dynamic buffer evolution: (a) anchor counts by source; (b) source composition along the quality scale.
Init.
K0
AE2
AH-v2
WB-v2
Writing
Avg.
Tokens (M)
Calls (K)
Nectar
2
80.99
30.00
46.73
66.69
56.10
79.3M
36.4K
4
82.80
30.20
48.79
67.27
57.27
88.8M
36.9K
7
81.06
29.60
48.02
66.92
56.40
114.1M
36.8K
Bootstrap
2
80.62
27.90
47.55
66.79
55.72
66.5M
28.0K
4
79.50
28.40
47.31
66.15
55.34
76.5M
28.4K
7
79.50
33.50
48.58
67.39
57.24
92.4M
28.7K
Table 4: Effect of the initial anchor count.
Variant
AE2
AH-v2
WB-v2
Writing
Avg.
Tokens (M)
Calls (K)
RankBuffer
82.80
30.20
48.79
67.27
57.27
88.8M
36.9K
Coarse-only
78.30
27.60
45.56
65.89
54.34
74.1M
31.0K
Table 5: Coarse-only ablation.
Initialization
Variant
AlpacaEval 2
Nectar
RankBuffer
82.80
Order-only
51.93
Bootstrap
RankBuffer
79.50
Order-only
54.53
Table 6: Order-only ablation.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Training
Queries per training step
20
Rollouts per query ( G )
8
Policy learning rate
1×10−6
Maximum prompt length
4,096 tokens
Maximum response length
4,096 tokens
Appendix
Table 7: Main implementation settings.
Method
AE2
AH-v2
WB-v2
Writing
Avg.
Tokens
Calls
GRPO
72.73
27.50
46.57
65.42
53.06
25,035,966
31,374
GDPO
64.35
20.10
45.14
64.94
48.63
86,937,684
144,236
Dr. GRPO
56.15
18.90
44.32
63.37
45.60
21,936,802
31,649
DAPO
74.22
26.60
46.83
66.13
53.44
24,078,025
31,231
GPG
53.04
16.90
45.78
64.98
45.18
20,915,602
31,619
Tournament-GRPO
80.75
29.50
48.87
68.16
56.82
157,313,284
84,000
Appendix
Table 8: Complete effectiveness and judge-cost comparison. An asterisk (*) denotes static buffers.
School of Computer Science and Engineering, Northeastern University, Shenyang, China · 2NiuTrans Research, Shenyang, China · 3Independent Researcher, Beijing, China +1