Generative reward models (GRMs) are important for LLM optimization. Unlike scalar reward models, GRMs generate natural-language critiques alongside preference judgments, providing finer-grained evaluation signals. Their effectiveness depends heavily on critique reliability. However, existing GRM training typically uses final preference correctness as outcome supervision. Because the preference outcome space is highly constrained, unreliable critiques can still yield correct outcomes and thus be reinforced. Recent work leverages human critiques for process supervision, but such critiques are scarce and are often reduced to scalar rewards, leaving their fine-grained evaluative information underutilized. We argue that evaluative criteria learned from human critiques can be generalized to broader outcome-only preference data. To this end, we propose \textbf{EnGRICH}, a GRM training framework that pairs the GRM with a training-time MetaCritic learned from a small set of human critiques. MetaCritic constructs response-specific rubrics and uses them to evaluate the evidence coverage and correctness of generated critiques. The resulting signals provide both process rewards for fine-grained credit assignment and structured guidance for exploring better critiques. During GRM training, MetaCritic is further optimized to generalize human-grounded evaluative criteria to outcome-only data. At inference, the trained GRM operates independently. Experiments across seven reward-model benchmarks show that EnGRICH consistently improves over competitive baselines, while further analyses validate the effectiveness of its core mechanisms.
Figures & tables
Figure 1: Illustration of EnGRICH.
Model / Method
HS3
RB2
RM-B
SCAN
HREF
LitBench
WQ
Overall
Open-Weight Reward Models
GRAM-Qwen3-4B
74.17
66.01
66.87
74.44
60.78
62.24
55.03
65.65
GRAM-R 2 -3B
76.37
50.94
64.18
69.97
63.19
62.64
59.16
63.78
CE-RM-4B
73.44
70.76
79.50
68.85
59.66
57.42
52.78
66.06
UltraRM-13B
74.60
53.07
57.20
74.31
58.26
69.80
59.06
63.76
RM-R1-14B
76.42
65.25
81.11
77.32
64.79
65.10
58.86
69.84
Table 1: Results across seven reward-model benchmarks. Bold and underlined indicate the best and second-best controlled results within each backbone; Overall is the unweighted mean.
Figure 2: Critique quality evaluation. All compared methods use Qwen3-4B as the GRM backbone. Left: CC-F1, OPI, and average assessment length on HelpSteer3 validation set; higher CC-F1 and lower OPI indicate better critique quality. Right: response-revision gains over direct-answer baselines.
Benchmark Performance
Critique Quality
Method
HS3
RB2
RM-B
SCAN
HREF
LitBench
WQ
CC-F1
OPI
EnGRICH
77.71
59.85
74.66
78.27
66.91
69.58
65.24
19.36
17.14
w/o Process Reward
76.74
56.90
72.60
78.43
63.15
67.16
62.07
14.96
25.81
w/o Guided Exploration
77.31
55.70
74.10
77.64
66.94
69.15
64.22
18.22
19.92
w/o Rubric Generalization
76.98
53.95
71.60
75.88
61.79
66.05
60.49
16.37
23.02
w/o DO Process Supervision
77.15
54.65
72.20
76.84
62.71
66.57
61.41
17.44
22.14
Table 2: Ablation study of EnGRICH on Gemma3-12B.
Figure 3: Scaling human critique supervision with Qwen3-4B as backbone. We report benchmark performance (left), critique quality (middle), and MetaCritic quality (right) as human-critique coverage increases. Initial denotes MetaCritic before online GRM training (SFT-initialized when human critiques are available), and Final denotes MetaCritic after GRM training.
Policy Setting
IFEval P-Strict
AlpacaEval 2 LC
Arena-Hard v2 CW
Initial policy
84.84
39.63
15.24
Outcome-GRPO
83.73
40.54
19.19
RC-GRPO
83.55
40.21
18.87
RM-NLHF
84.10
40.43
19.74
EnGRICH
87.06
42.80
21.43
Table 3: Downstream performance of Qwen3-4B-Instruct-2507 policies trained with Gemma3-12B as the GRM backbone.
Figure 4: Training and inference efficiency of RL-based GRMs on Gemma3-12B. Aux. Update denotes MetaRM updates in RM-NLHF and Mrub updates in EnGRICH. Inference time and output length are averaged per sample across seven benchmarks.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Training prompt for the GRM’s initial assessment. The output separates critique text from the final preference decision.
Figure 6: Rubric-generation prompt shared across the two training subsets. The rubric generator receives only the task and candidate responses; human critiques, gold preference labels, and sampled GRM critiques are not exposed to this generation call.
Figure 7: Prompt for the frozen critique verifier Mver . Evaluation points retain the rubric generator’s numbered natural-language format. The displayed output lines illustrate the syntax; the actual number of labels equals the number of supplied keypoints.
Figure 8: Diagnosis-guided revision prompt used for all-failed groups. The GRM uses the keypoint-level diagnoses as guidance, verifies them against the original inputs, and then produces a revised critique and preference without access to the gold preference or an explicit correctness signal.
Domain
Dataset
Pairs
Share
Source
Human-critique subset DH
Mixed
HelpSteer3-Preference
12,000
20.00%
RM-NLHF
Outcome-only subset DO
General
HelpSteer2
3,000
5.00%
Skywork-Reward- [-0.4pt] Preference-80K-v0.2
General
Magpie-Pro-Llama-3.1
8,000
13.33%
General
Magpie-Pro
1,000
1.67%
Appendix
Table 4: Composition of the 60K training mixture. DH includes human-derived critique supervision, whereas DO contains preference supervision only.
Benchmark
Examples
Comparisons
Metric
Source
HelpSteer3 (HS3)
1,287
1,287
Pairwise accuracy
RM-NLHF
RewardBench 2 (RB2)
1,763
5,289
Mean domain accuracy
RewardBench-2
RM-Bench (RM-B)
1,327
11,943
Mean domain accuracy
RM-Bench
SCAN-HPD
626
626
Pairwise accuracy
SCAN-Dataset
HREF-HA
1,452
1,452
LOO human agreement
HREF-Preference
LitBench
2,480
4,960
Order-averaged accuracy
LitBench-Test
Appendix
Table 5: Evaluation suites used in our experiments. Counts correspond to the benchmark manifests before input-length filtering. Examples denotes benchmark records and comparisons denotes pairwise GRM predictions.
Setting
Qwen3-4B
Gemma3-12B
Judge LoRA (r,α)
(128,256)
(128,256)
Judge adapter parameters
264.24M
523.76M
Each MetaCritic LoRA (r,α)
(64,128)
(64,128)
Each MetaCritic adapter parameters
132.12M
261.88M
Global Judge batch size
64
64
Judge prompt / response cap
4,096 / 2,048
4,096 / 2,048
Appendix
Table 6: Principal adapter sizes and training budgets. Adapter sizes count LoRA parameters only. The verifier adapter is frozen during online training. Online budgets refer to one pass over the tokenizer-filtered preference mixture.
Figure 9: Brief critique-generation prompt shared by the compared GRMs. Each input contains the question and the same two candidate responses.
Figure 10: Prompts for CC-F1. Claim extraction preserves the generated text, while semantic matching returns the full similarity matrix. The final one-to-one assignment is computed locally.
Figure 11: Core-argument scoring prompt from RM-NLHF ( Wang et al., 2026c ) . The resulting F1 is thresholded at 0.5 when computing OPI.
Benchmark
Domain
Examples
Metric
Source
MATH-500
Math
500
Answer accuracy
MATH-500
HumanEval+
Code
164
Pass@1
EvalPlus
Arena-Hard-v2.0
Mixed
500
Tie-adjusted win rate
Arena-Hard-v2.0
Total
1,164
–
Appendix
Table 7: Downstream datasets for critique-guided response refinement. Counts denote questions evaluated under each critique-model/editor combination.
Figure 12: RM-NLHF editing instruction with a complete-answer return convention. The editor receives both responses selected by the GRM and its critique, and returns one final response.
Figure 13: Arena-Hard judge prompt for comparing a response against the fixed reference. The candidate and reference are evaluated in both orders; the same prompt is used for direct answers and final Feedback-Edit responses.
Figure 14: Critique quality evaluation on Qwen3-4B and Gemma-3-12B. Direct evaluation reports CC-F1, OPI, and average assessment length on the HelpSteer3 validation set, while downstream evaluation reports response-revision gains over direct-answer baselines.
Benchmark Performance
Critique Quality
Method
HS3
RB2
RM-B
SCAN
HREF
LitBench
WQ
CC-F1
OPI
Qwen3-4B
EnGRICH
77.06
65.42
75.12
77.48
62.60
65.69
61.15
18.18
28.09
w/o Process Reward
75.95
62.10
73.74
76.04
58.61
64.23
57.94
14.46
34.17
w/o Guided Exploration
76.82
60.95
74.58
77.32
61.90
65.73
60.54
17.31
30.48
w/o Rubric Generalization
77.14
58.90
72.90
75.72
57.95
63.71
57.38
15.65
33.85
Appendix
Table 8: Ablation results on Qwen3-4B and Gemma3-12B. Bold denotes the best result within each backbone.
Figure 15: Scaling human critique supervision on Qwen3-4B and Gemma3-12B. We report benchmark performance (left), critique quality (middle), and MetaCritic quality (right) as human-critique coverage increases. Initial denotes MetaCritic before online GRM training (SFT-initialized when human critiques are available), and Final denotes MetaCritic after GRM training.
Figure 16: Training and inference efficiency of RL-based GRMs on Qwen3-4B and Gemma3-12B. Aux. Update denotes MetaRM updates in RM-NLHF and Mrub updates in EnGRICH. Inference time and output length are averaged per sample across seven benchmarks.
Reliable reward and preference signals are critical for evaluating and optimizing large language models on open-ended tasks. Rubric-based judges offer a transparent way to decompose such judgments into explicit evaluation criteria, but existing annotation-free rubric generators typically rely on a single generic evaluator. As a result, they may overlook important dimensions of human preference, a failure mode we term dimensional blind spots. To address this limitation, we propose Multi-Role Rubric Generation (MRRG), a training-free and reference-free framework that elicits evaluation criteria from multiple complementary roles and consolidates them into an auditable rubric-based scorer. This scorer can be used both to validate pairwise preferences and to provide rewards for GRPO-style Reinforcement Learning with Verifiable Rewards (RLVR). Experiments on preference validation benchmarks show that MRRG consistently outperforms single-role rubric generation baselines across multiple backbone models. Further RLVR experiments demonstrate that MRRG yields a stronger reward signal for improving open-ended generation.
Dazhi Fu, Jiuding Yang, Yiwen Guo +1
School of Data Science, The Chinese University of Hong Kong, Shenzhen, China · LIGHTSPEED · Independent Researcher
Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models. However, despite their strong capabilities in response ranking, generative reward models have not realized their potential in reinforcement learning (RL). Our analysis reveals that this limitation arises from a mismatch between the comparative nature of generative reward modeling and the scalar scoring paradigm adopted by existing RL algorithms. To bridge this gap, we propose a Ranking-based Reward Construction (RRC) approach, which enables generative reward models to provide more effective RL learning signals by deriving rewards from relative preference rankings. RRC introduces two complementary strategies: self-competitive ranking, which exploits comparisons among sampled responses, and anchor-guided ranking, which enables scalable ranking-based reward construction with a small set of reference responses. Experiments across open-ended chat and reasoning benchmarks demonstrate that RRC substantially improves RL training with generative reward models, achieving consistent gains over existing reward construction approaches. Our code can be found at https://github.com/wangclnlp/RRC.
Chenglong Wang, Ziming Zhu, Yifu Huo +9
School of Computer Science and Engineering, Northeastern University, Shenyang, China · 2NiuTrans Research, Shenyang, China · 3Independent Researcher, Beijing, China +1
Reinforcement Learning from Human Feedback has become the standard paradigm for language model alignment, where reward models directly determine alignment effectiveness. In this work, we focus on how to evaluate the generalizability of reward models. By "generalizability", we mean the ability of RMs to correctly rank responses to align with diverse user preferences. However, existing reward model benchmarks are typically designed around a universal preference, failing to assess this generalization. To address this critical gap, we introduce RMGAP, a benchmark comprising 1,097 instances across Chat, Writing, Reasoning, and Safety domains. Since different users exhibit diverse preferences for the same task, we first generate four distinct responses with different linguistic profiles for each collected prompt. However, the original prompt set lacks the specificity to convey different preferences. We therefore construct tailored prompts by contrasting these candidates and designing scenarios in which one response becomes the uniquely appropriate choice. Moreover, we observe that users often express the same preference using different phrasings, and thus extend each prompt with two paraphrased variants. Our evaluation of 24 state-of-the-art RMs reveals their substantial limitations: even the best RM achieves only 49.27% Best-of-N accuracy, highlighting considerable room for improvement in reward model generalization. Related data and code are available at https://github.com/nanzhi84/RMGAP.
Yangyang Zhou, Yi-Chen Li
Beijing University of Posts and Telecommunications · National Key Laboratory for Novel Software Technology, Nanjing University, China · School of Artificial Intelligence, Nanjing University, China