Open-ended query generation lacks standard answers, thus necessitating an effective reward mechanism. Pointwise scoring rubrics provide limited information about the relative quality of sample answers under the same prompt; merging multiple rubric judgments into a single score may also mask the differences between these answers. We propose MatrixReward, which constructs rewards from a rollout-by-rubric win-rate matrix obtained by comparing every pair of sampled responses under each rubric. The spread of each matrix column captures how strongly that rubric distinguishes the current rollouts, while correlations between columns reveal rubric repetition; together, these statistics yield data-dependent rubric weights. We combine these weights with the prior weights of rubrics. After column normalization and weighting, the observed per-rubric maxima and minima define positive and negative ideal profiles. Each rollout's distances to these two ideals determine its relative-closeness quality reward. Evaluated using Qwen3-8B on four open-ended query-answering benchmarks, MatrixReward achieves an average score of 63.02, outperforming the strongest baseline by approximately 2.0%. These results support the idea that matrices derived from relative comparisons can be used to construct rewards more reasonably for open-ended generative reinforcement learning.
Figures & tables
Figure 1: Comparison of reward construction for the same rollout group.
Figure 2: Overview of the MatrixReward framework.
Method
AE2 ↑
AH2 ↑
Wild ↑
Writing ↑
Avg ↑
MatrixReward
85.35
41.23
61.83
63.66
63.02
MatrixReward (implicit)
83.45
40.90
61.09
63.47
62.23
Pointwise + MatrixReward
72.33
36.91
60.40
61.81
57.86
Pairwise baselines
RRC-SCR
83.60
39.98
60.59
63.00
61.79
ArenaRL
76.04
34.82
61.87
61.73
58.62
Table 1: Comparison with multiple types of baselines.
Weight rule
AE2 ↑
AH2 ↑
Wild ↑
Writing ↑
Avg ↑
MatrixReward
85.35
41.23
61.83
63.66
63.02
Prior weights
81.25
38.44
59.55
61.73
60.24
Entropy
81.61
40.75
60.53
62.95
61.46
SVD
82.12
41.22
61.83
63.45
62.16
Table 2: Rubric-wise matrix weight variants.
Variant
AE2 ↑
AH2 ↑
Wild ↑
Writing ↑
Avg ↑
MatrixReward
85.35
41.23
61.83
63.66
63.02
Weighted aggregation
78.02
38.54
60.32
62.45
59.83
Theoretical ideals
81.79
40.02
59.74
61.67
60.81
Mahalanobis distance
80.38
40.41
60.66
62.73
61.05
Table 3: MatrixReward reward-aggregation and distance comparisons.
Figure 3: Reward scores under alternative ideal profiles and distance metrics. (a) Empirical versus theoretical ideals. (b) Euclidean distance in MatrixReward versus Mahalanobis distance with the same empirical ideals.
Figure 4: A response group with eight rollouts and three rubrics. (a) Column-normalized profiles and observed ideals. (b) Weighted aggregation and MatrixReward quality rewards. The rubric weights are (wc1,wc2,wc3)=(0.452,0.245,0.304) .
Fusion rule
AE2 ↑
AH2 ↑
Wild ↑
Writing ↑
Avg ↑
MatrixReward
85.35
41.23
61.83
63.66
63.02
Additive ( α=0.1 )
81.71
41.40
60.31
62.46
61.47
Additive ( α=0.3 )
82.42
40.40
60.73
62.85
61.60
Additive ( α=0.5 )
83.95
40.51
61.23
62.45
62.04
Multiplicative
80.51
36.52
60.73
62.15
59.98
Table 4: MatrixReward weight-fusion ablation.
Method
Avg ↑
Calls/group
Total calls (M)
Total tokens (B)
Tokens/call
MatrixReward
63.02
112
0.86
2.00
2,321
MatrixReward (implicit)
62.23
28
0.22
0.52
2,408
RRC-SCR
61.79
28
0.21
0.58
2,676
SAW
57.42
8
0.06
0.07
1,070
DAPO
56.09
8
0.06
0.07
1,070
Table 5: Benchmark average and training judge cost.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Data and model
Training queries
2,560 (UltraFeedback)
Query-specific rubrics per query
3–5
Rubric importance weights
1 – 3
Policy backbone
Qwen3-8B
Training judge
Qwen3.6-35B-A3B
Appendix
Table 6: Main implementation hyperparameters.
Method
AE2 ↑
AH2 ↑
Wild ↑
Writing ↑
Avg ↑
MatrixReward
70.60
20.89
54.16
53.70
49.84
MatrixReward (implicit)
68.72
19.58
53.38
51.89
48.39
RRC-SCR
68.48
19.66
52.29
51.74
48.04
ArenaRL
63.42
18.03
53.17
53.54
47.04
Tournament-GRPO
62.64
17.95
53.46
52.28
46.58
Pairwise Rank
64.27
19.05
54.60
52.77
47.67
Appendix
Table 7: Qwen2.5-7B-Instruct benchmark results.
Rubric
Description
Points
c1
The answer is an original poem, not prose, advice, explanation, unsolicited criticism, or moralizing.
3
c2
It clearly expresses the requested contrast: the girl does not cause fleeting “butterflies” because they leave, while the speaker’s feelings for her will not leave and are here to stay.
3
c3
It uses effective poetic craft—imagery, metaphor, rhythm, and/or rhyme—to make the sentiment lyrical rather than generic or awkward.
2
c4
It maintains a sincere, affectionate, first-person romantic tone appropriate to the requested sentiment.
2
Appendix
Table 8: Query-specific rubrics for the poetry request.
Response
c1
c2
c3
c4
y1
0.286
0.643
0.143
0.357
y2
0.286
0.143
0.214
0.071
y3
0.286
0.000
0.286
0.071
y4
0.357
0.500
0.643
0.786
y5
0.893
0.643
0.929
0.786
y6
0.643
0.857
0.429
0.714
Appendix
Table 9: Rubric-wise pairwise win rates M for the eight sampled poems.
Quantity
c1
c2
c3
c4
Column deviation σk
0.3869
0.3436
0.3636
0.3884
Conflict hk
0.6694
0.9218
0.7780
0.5521
Matrix-derived weight ak
0.2413
0.2952
0.2637
0.1998
Prior weight pk
0.3000
0.3000
0.2000
0.2000
Combined weight wk
0.2907
0.2992
0.2101
0.2000
Column norm Lk
1.5625
1.6413
1.6288
1.6568
Appendix
Table 10: Weight extraction, combination, and ideal profiles in the poetry group.
Response
Di+
Di−
riqual
ri
Ai
y1
0.1679
0.1222
0.4211
1.4211
-0.2418
y2
0.2173
0.0276
0.1128
1.1128
-1.2082
y3
0.2303
0.0184
0.0741
1.0741
-1.3297
y4
0.1249
0.1417
0.5314
1.5314
0.1039
y5
0.0400
0.2102
0.8401
1.8401
1.0714
y6
0.0814
0.1903
0.7005
1.7005
0.6338
Appendix
Table 11: Ideal-point distances, rewards, and group-normalized advantages for the poetry group.