cs.AIOct 4, 2026

EnGRICH: Enhancing Generative Reward Modeling with Critiques from Humans

Authors: Xuancheng Li, Beining Wang, Haitao Li, Heng Wang, Yujia Zhou, Qingyi Pan, Blaze Chen, Yiqun Liu, +2 more

Organizations: Department of Computer Science and Technology, Tsinghua University, Beijing, China · Tencent, Beijing, China

Abstract

Generative reward models (GRMs) are important for LLM optimization. Unlike scalar reward models, GRMs generate natural-language critiques alongside preference judgments, providing finer-grained evaluation signals. Their effectiveness depends heavily on critique reliability. However, existing GRM training typically uses final preference correctness as outcome supervision. Because the preference outcome space is highly constrained, unreliable critiques can still yield correct outcomes and thus be reinforced. Recent work leverages human critiques for process supervision, but such critiques are scarce and are often reduced to scalar rewards, leaving their fine-grained evaluative information underutilized. We argue that evaluative criteria learned from human critiques can be generalized to broader outcome-only preference data. To this end, we propose \textbf{EnGRICH}, a GRM training framework that pairs the GRM with a training-time MetaCritic learned from a small set of human critiques. MetaCritic constructs response-specific rubrics and uses them to evaluate the evidence coverage and correctness of generated critiques. The resulting signals provide both process rewards for fine-grained credit assignment and structured guidance for exploring better critiques. During GRM training, MetaCritic is further optimized to generalize human-grounded evaluative criteria to outcome-only data. At inference, the trained GRM operates independently. Experiments across seven reward-model benchmarks show that EnGRICH consistently improves over competitive baselines, while further analyses validate the effectiveness of its core mechanisms.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Many Voices, One Reward: Multi-Role Rubric Generation for LLM Judging and Reward Modeling

    Jul 2, 2026Dazhi Fu, Jiuding Yang, Yiwen Guo +1Task-Specific RubricsProgress Reward Modeling

  2. RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction

    Aug 6, 2026Chenglong Wang, Ziming Zhu, Yifu Huo +9Large Language Model Reinforcement Learning

  3. RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences

    May 3, 2026Yangyang Zhou, Yi-Chen LiProgress Reward ModelingLarge Language Model Alignment