Generative reward models (GRMs) are important for LLM optimization. Unlike scalar reward models, GRMs generate natural-language critiques alongside preference judgments, providing finer-grained evaluation signals. Their effectiveness depends heavily on critique reliability. However, existing GRM training typically uses final preference correctness as outcome supervision. Because the preference outcome space is highly constrained, unreliable critiques can still yield correct outcomes and thus be reinforced. Recent work leverages human critiques for process supervision, but such critiques are scarce and are often reduced to scalar rewards, leaving their fine-grained evaluative information underutilized. We argue that evaluative criteria learned from human critiques can be generalized to broader outcome-only preference data. To this end, we propose \textbf{EnGRICH}, a GRM training framework that pairs the GRM with a training-time MetaCritic learned from a small set of human critiques. MetaCritic constructs response-specific rubrics and uses them to evaluate the evidence coverage and correctness of generated critiques. The resulting signals provide both process rewards for fine-grained credit assignment and structured guidance for exploring better critiques. During GRM training, MetaCritic is further optimized to generalize human-grounded evaluative criteria to outcome-only data. At inference, the trained GRM operates independently. Experiments across seven reward-model benchmarks show that EnGRICH consistently improves over competitive baselines, while further analyses validate the effectiveness of its core mechanisms.
Figures & tables
Figure 1: Illustration of EnGRICH.
Model / Method
HS3
RB2
RM-B
SCAN
HREF
LitBench
WQ
Overall
Open-Weight Reward Models
GRAM-Qwen3-4B
74.17
66.01
66.87
74.44
60.78
62.24
55.03
65.65
GRAM-R 2 -3B
76.37
50.94
64.18
69.97
63.19
62.64
59.16
63.78
CE-RM-4B
73.44
70.76
79.50
68.85
59.66
57.42
52.78
66.06
UltraRM-13B
74.60
53.07
57.20
74.31
58.26
69.80
59.06
63.76
RM-R1-14B
76.42
65.25
81.11
77.32
64.79
65.10
58.86
69.84
Table 1: Results across seven reward-model benchmarks. Bold and underlined indicate the best and second-best controlled results within each backbone; Overall is the unweighted mean.
Figure 2: Critique quality evaluation. All compared methods use Qwen3-4B as the GRM backbone. Left: CC-F1, OPI, and average assessment length on HelpSteer3 validation set; higher CC-F1 and lower OPI indicate better critique quality. Right: response-revision gains over direct-answer baselines.
Benchmark Performance
Critique Quality
Method
HS3
RB2
RM-B
SCAN
HREF
LitBench
WQ
CC-F1
OPI
EnGRICH
77.71
59.85
74.66
78.27
66.91
69.58
65.24
19.36
17.14
w/o Process Reward
76.74
56.90
72.60
78.43
63.15
67.16
62.07
14.96
25.81
w/o Guided Exploration
77.31
55.70
74.10
77.64
66.94
69.15
64.22
18.22
19.92
w/o Rubric Generalization
76.98
53.95
71.60
75.88
61.79
66.05
60.49
16.37
23.02
w/o DO Process Supervision
77.15
54.65
72.20
76.84
62.71
66.57
61.41
17.44
22.14
Table 2: Ablation study of EnGRICH on Gemma3-12B.
Figure 3: Scaling human critique supervision with Qwen3-4B as backbone. We report benchmark performance (left), critique quality (middle), and MetaCritic quality (right) as human-critique coverage increases. Initial denotes MetaCritic before online GRM training (SFT-initialized when human critiques are available), and Final denotes MetaCritic after GRM training.
Policy Setting
IFEval P-Strict
AlpacaEval 2 LC
Arena-Hard v2 CW
Initial policy
84.84
39.63
15.24
Outcome-GRPO
83.73
40.54
19.19
RC-GRPO
83.55
40.21
18.87
RM-NLHF
84.10
40.43
19.74
EnGRICH
87.06
42.80
21.43
Table 3: Downstream performance of Qwen3-4B-Instruct-2507 policies trained with Gemma3-12B as the GRM backbone.
Figure 4: Training and inference efficiency of RL-based GRMs on Gemma3-12B. Aux. Update denotes MetaRM updates in RM-NLHF and Mrub updates in EnGRICH. Inference time and output length are averaged per sample across seven benchmarks.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Training prompt for the GRM’s initial assessment. The output separates critique text from the final preference decision.
Figure 6: Rubric-generation prompt shared across the two training subsets. The rubric generator receives only the task and candidate responses; human critiques, gold preference labels, and sampled GRM critiques are not exposed to this generation call.
Figure 7: Prompt for the frozen critique verifier Mver . Evaluation points retain the rubric generator’s numbered natural-language format. The displayed output lines illustrate the syntax; the actual number of labels equals the number of supplied keypoints.
Figure 8: Diagnosis-guided revision prompt used for all-failed groups. The GRM uses the keypoint-level diagnoses as guidance, verifies them against the original inputs, and then produces a revised critique and preference without access to the gold preference or an explicit correctness signal.
Domain
Dataset
Pairs
Share
Source
Human-critique subset DH
Mixed
HelpSteer3-Preference
12,000
20.00%
RM-NLHF
Outcome-only subset DO
General
HelpSteer2
3,000
5.00%
Skywork-Reward- [-0.4pt] Preference-80K-v0.2
General
Magpie-Pro-Llama-3.1
8,000
13.33%
General
Magpie-Pro
1,000
1.67%
Appendix
Table 4: Composition of the 60K training mixture. DH includes human-derived critique supervision, whereas DO contains preference supervision only.
Benchmark
Examples
Comparisons
Metric
Source
HelpSteer3 (HS3)
1,287
1,287
Pairwise accuracy
RM-NLHF
RewardBench 2 (RB2)
1,763
5,289
Mean domain accuracy
RewardBench-2
RM-Bench (RM-B)
1,327
11,943
Mean domain accuracy
RM-Bench
SCAN-HPD
626
626
Pairwise accuracy
SCAN-Dataset
HREF-HA
1,452
1,452
LOO human agreement
HREF-Preference
LitBench
2,480
4,960
Order-averaged accuracy
LitBench-Test
Appendix
Table 5: Evaluation suites used in our experiments. Counts correspond to the benchmark manifests before input-length filtering. Examples denotes benchmark records and comparisons denotes pairwise GRM predictions.
Setting
Qwen3-4B
Gemma3-12B
Judge LoRA (r,α)
(128,256)
(128,256)
Judge adapter parameters
264.24M
523.76M
Each MetaCritic LoRA (r,α)
(64,128)
(64,128)
Each MetaCritic adapter parameters
132.12M
261.88M
Global Judge batch size
64
64
Judge prompt / response cap
4,096 / 2,048
4,096 / 2,048
Appendix
Table 6: Principal adapter sizes and training budgets. Adapter sizes count LoRA parameters only. The verifier adapter is frozen during online training. Online budgets refer to one pass over the tokenizer-filtered preference mixture.
Figure 9: Brief critique-generation prompt shared by the compared GRMs. Each input contains the question and the same two candidate responses.
Figure 10: Prompts for CC-F1. Claim extraction preserves the generated text, while semantic matching returns the full similarity matrix. The final one-to-one assignment is computed locally.
Figure 11: Core-argument scoring prompt from RM-NLHF ( Wang et al., 2026c ) . The resulting F1 is thresholded at 0.5 when computing OPI.
Benchmark
Domain
Examples
Metric
Source
MATH-500
Math
500
Answer accuracy
MATH-500
HumanEval+
Code
164
Pass@1
EvalPlus
Arena-Hard-v2.0
Mixed
500
Tie-adjusted win rate
Arena-Hard-v2.0
Total
1,164
–
Appendix
Table 7: Downstream datasets for critique-guided response refinement. Counts denote questions evaluated under each critique-model/editor combination.
Figure 12: RM-NLHF editing instruction with a complete-answer return convention. The editor receives both responses selected by the GRM and its critique, and returns one final response.
Figure 13: Arena-Hard judge prompt for comparing a response against the fixed reference. The candidate and reference are evaluated in both orders; the same prompt is used for direct answers and final Feedback-Edit responses.
Figure 14: Critique quality evaluation on Qwen3-4B and Gemma-3-12B. Direct evaluation reports CC-F1, OPI, and average assessment length on the HelpSteer3 validation set, while downstream evaluation reports response-revision gains over direct-answer baselines.
Benchmark Performance
Critique Quality
Method
HS3
RB2
RM-B
SCAN
HREF
LitBench
WQ
CC-F1
OPI
Qwen3-4B
EnGRICH
77.06
65.42
75.12
77.48
62.60
65.69
61.15
18.18
28.09
w/o Process Reward
75.95
62.10
73.74
76.04
58.61
64.23
57.94
14.46
34.17
w/o Guided Exploration
76.82
60.95
74.58
77.32
61.90
65.73
60.54
17.31
30.48
w/o Rubric Generalization
77.14
58.90
72.90
75.72
57.95
63.71
57.38
15.65
33.85
Appendix
Table 8: Ablation results on Qwen3-4B and Gemma3-12B. Bold denotes the best result within each backbone.
Figure 15: Scaling human critique supervision on Qwen3-4B and Gemma3-12B. We report benchmark performance (left), critique quality (middle), and MetaCritic quality (right) as human-critique coverage increases. Initial denotes MetaCritic before online GRM training (SFT-initialized when human critiques are available), and Final denotes MetaCritic after GRM training.
Figure 16: Training and inference efficiency of RL-based GRMs on Qwen3-4B and Gemma3-12B. Aux. Update denotes MetaRM updates in RM-NLHF and Mrub updates in EnGRICH. Inference time and output length are averaged per sample across seven benchmarks.
School of Computer Science and Engineering, Northeastern University, Shenyang, China · 2NiuTrans Research, Shenyang, China · 3Independent Researcher, Beijing, China +1
Beijing University of Posts and Telecommunications · National Key Laboratory for Novel Software Technology, Nanjing University, China · School of Artificial Intelligence, Nanjing University, China