cs.LGSep 30, 2026

Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics

Authors: Maoqi Liu, Junwei He, Bowen Zhang, Feiran Li, Wentao Ma, Rongyi Lin, Shuhan Zhong, Quan Fang

Organizations: Beijing University of Posts and Telecommunications · ByteDance

Abstract

Rubric-based reinforcement learning (Rubric-RL) trains language models where no verifier exists. A judge checks each criterion of a rubric, and the verdicts are aggregated into a reward, most often by a weighted sum. We show that this additive aggregation is the weak point. Under a sum, criteria compensate for one another: a policy that misses the one decision that matters can buy the points back with advice nobody asked for. On clinical consultation, such a policy scores higher and answers worse. Rubric coverage rises while appropriateness on held-out physician criteria falls below the untrained model. The medical criteria are not to blame. Grouped so that they must hold together, the same criteria, unchanged to the word, recover a third of the loss; shorter answers recover almost none. We therefore propose Protocol-level Rubrics (ProRubric), which keeps what the criteria ask for and changes how they are aggregated. It groups a checklist into a few protocol-level dimensions. A dimension counts only when all of its criteria hold and its failure clause does not fire. The grouping is done once, offline, and leaves the optimizer unchanged. ProRubric raises appropriateness by 10.8 points without losing coverage and has the best seven-benchmark average at both scales. Reward validity is set not only by what a rubric verifies, but by how it aggregates. Code is available at https://github.com/Estrellajer/ProRubric

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL

    Aug 12, 2026Minglai Yang, Xinyu Guo, Utkarsh Tyagi +6Task-Specific RubricsRubrics

  2. Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR

    May 19, 2026Utkarsh Tyagi, Xingang Guo, MohammadHossein Rezaei +5Rubric-Based ScoringReward Functions

  3. Reward Hacking in Rubric-Based Reinforcement Learning

    May 12, 2026Anas Mahmoud, MohammadHossein Rezaei, Zihao Wang +3Verifiable Rewards