Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics
Organizations: Beijing University of Posts and Telecommunications · ByteDance
Abstract
Rubric-based reinforcement learning (Rubric-RL) trains language models where no verifier exists. A judge checks each criterion of a rubric, and the verdicts are aggregated into a reward, most often by a weighted sum. We show that this additive aggregation is the weak point. Under a sum, criteria compensate for one another: a policy that misses the one decision that matters can buy the points back with advice nobody asked for. On clinical consultation, such a policy scores higher and answers worse. Rubric coverage rises while appropriateness on held-out physician criteria falls below the untrained model. The medical criteria are not to blame. Grouped so that they must hold together, the same criteria, unchanged to the word, recover a third of the loss; shorter answers recover almost none. We therefore propose Protocol-level Rubrics (ProRubric), which keeps what the criteria ask for and changes how they are aggregated. It groups a checklist into a few protocol-level dimensions. A dimension counts only when all of its criteria hold and its failure clause does not fire. The grouping is done once, offline, and leaves the optimizer unchanged. ProRubric raises appropriateness by 10.8 points without losing coverage and has the best seven-benchmark average at both scales. Reward validity is set not only by what a rubric verifies, but by how it aggregates. Code is available at https://github.com/Estrellajer/ProRubric
Figures & tables
| (a) After Rubric-RL training: untrained answer preferred (%) | (b) Before training: reward change per edit | |||||
| Domain | DeepSeek-V4-Pro | GPT-5.6-luna | Edit to the answer | Rubric-RL | raw-AND | ProRubric |
| Medicine | 13.2 | 2.3 | Name an item | |||
| Science | 1.2 | 9.0 | Name the topic | |||
| Dialogue | 5.8 | 6.8 | Add a needless test | |||
| Writing | 0.0 | 1.6 | Drop the key advice | |||
| Method | Writing | Dialogue | Clinical Medicine | Scientific Reasoning | Overall | |||
|---|---|---|---|---|---|---|---|---|
| WritingBench | Creative-v3 | Arena-Hard | HealthBench | MedQA | GPQA | ResearchQA | Average | |
| Qwen3-4B | ||||||||
| Base | 49.2 | 27.9 | 12.9 | 26.1 | 65.2 | 41.1 | 51.8 | 39.2 |
| + SFT | 63.9 14.7 | 35.5 7.6 | 21.6 8.7 | 35.1 9.1 | 59.5 5.7 | 36.0 5.1 | 63.2 11.4 | 45.0 5.8 |
| + OPSD | 49.7 0.4 | 27.7 0.1 | 9.2 3.7 | 25.1 0.9 | 60.8 4.4 | 36.4 4.7 | 47.5 4.3 | 36.6 2.5 |
| + Rubric-RL | 64.6 15.3 | 35.0 7.2 | 17.3 4.4 | 31.9 5.8 | 62.6 2.6 | 44.7 3.6 | 62.9 11.1 | 45.6 6.4 |
| Reward | ||
|---|---|---|
| ProRubric variants | ||
| ProRubric | 36.8 2.7 | 35.0 1.4 |
| failure clauses | 34.5 1.1 | 34.9 0.3 |
| partial credit (Graded) | 36.3 1.8 | 35.1 0.9 |
| one dimension ( ) | 36.9 1.9 | 33.5 1.2 |
| Same criteria, other aggregation |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Rubric-RL / ProRubric | RuscaRL | OPSD | SFT |
|---|---|---|---|---|
| Training steps | 300 | 300 | 485 | 1,224 (3 epochs) |
| Prompts per batch | 64 | 512 | 128 | 64 |
| Responses per prompt | 8 | 1 | 1 | — |
| Mini-batch size | 32 | 256 | 32 | — |
| Learning rate | ||||
| Prompt limit | 4,096 | 6,144 | 4,096 | 20,000 |
| Metric | Emergency referrals | Global health | Communi- cation | Context seeking | Hedging | Health data | Complex responses | All criteria |
|---|---|---|---|---|---|---|---|---|
| Criteria ( ) | 6 | 3 | 4 | 4 | 9 | 4 | 4 | 34 |
| DeepSeek-V4-Pro | 0.65 | 0.69 | 0.63 | 0.61 | 0.61 | 0.71 | 0.45 | 0.62 |
| Physicians | 0.65 | 0.65 | 0.62 | 0.64 | 0.66 | 0.74 | 0.57 | 0.65 |
| Medicine | Science | Writing | Dialogue |
|---|---|---|---|
| HealthBench-consensus (3,140) | GPQA-Diamond (198) | WritingBench (1,000) | Arena-Hard v2 (748) |
| HealthBench-full, lite (4,392) | ResearchQA (702) | Creative-v3 (96) | |
| HealthBench-full, pro (4,138) | |||
| MedQA (1,273) |
| DeepSeek-V4-Pro | GPT-5.6-luna | |||
| Domain | W / T / L | Win rate (%) | W / T / L | Win rate (%) |
| Base vs. Rubric-RL | ||||
| Medicine | 136 / 93 / 71 | 65.7 13.2 | 229 / 55 / 16 | 93.6 2.3 |
| Science | 3 / 35 / 262 | 1.2 1.2 | 151 / 134 / 15 | 90.3 9.0 |
| Dialogue | 45 / 117 / 138 | 24.3 5.8 | 81 / 114 / 99 | 44.9 6.8 |
| Writing | 0 / 8 / 292 | 0.0 0.0 | 4 / 30 / 266 | 1.4 1.6 |
| Benchmark | Rubric-RL | ProRubric | Paired | ||
|---|---|---|---|---|---|
| per seed | mean | per seed | mean | ||
| Writing | |||||
| WritingBench | 66.3 / 63.6 / 63.7 | 64.6 1.5 | 67.9 / 66.8 / 66.5 | 67.1 0.7 | +2.50 0.85 |
| Creative-v3 | 36.5 / 34.1 / 34.5 | 35.0 1.3 | 39.1 / 36.8 / 37.4 | 37.7 1.2 | +2.73 0.17 |
| Dialogue | |||||
| Arena-Hard | 18.2 / 16.1 / 17.4 | 17.3 1.1 | 18.4 / 17.5 / 19.8 | 18.6 1.1 | +1.29 1.07 |
| Method | per seed | per seed | per seed |
|---|---|---|---|
| Baselines | |||
| Rubric-RL | 24.3 / 28.4 / 25.1 2.2 | 69.9 / 71.3 / 68.3 1.5 | 55.7 / 56.1 / 55.4 0.4 |
| RuscaRL | 24.6 / 25.8 / 29.5 2.6 | 66.4 / 70.3 / 71.9 2.8 | 54.2 / 56.4 / 55.6 1.1 |
| OPSD | 40.7 / 42.9 / 41.4 1.1 | 64.8 / 67.9 / 66.3 1.6 | 35.4 / 36.9 / 36.5 0.8 |
| Aggregation changed | |||
| Implicit | 36.7 / 36.6 / 36.8 0.1 | 78.0 / 77.5 / 77.7 0.2 | 56.4 / 56.1 / 55.6 0.4 |
| Configuration | Reward structure | Score | Len. | |||||
|---|---|---|---|---|---|---|---|---|
| grouped | appr. crit. | |||||||
| Baselines | ||||||||
| Base | – | – | 38.1 | 26.6 | 78.0 | 54.1 | 3.0k | 1 |
| Rubric-RL | 55.8 0.4 | 32.1 | 69.8 1.5 | 26.0 2.2 | 11.7k | 3 | ||
| RuscaRL | 55.4 1.1 | 32.1 | 69.5 2.8 | 26.7 2.6 | 11.9k | 3 | ||
| OPSD | – | – | 36.3 0.8 | 25.6 | 66.3 1.6 | 41.7 1.1 | 4.7k | 3 |
| Baselines | Aggregation changed | + Appr. crit. | KL penalty | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Metric | Base | Rubric-RL | Implicit | raw-AND | Graded | ProRubric | Rubric-RL | ProRubric | Rubric-RL | Rubric-RL | ProRubric |
| + appr. | + appr. | ||||||||||
| Score | 33.7 | 14.4 | 21.6 | 22.4 | 20.1 | 20.3 | 18.0 | 22.6 | 20.2 | 29.2 | 28.5 |
| Items | 3,655 | 3,670 | 3,652 | 3,669 | 3,670 | 3,671 | 3,668 | 3,669 | 3,652 | 3,654 | 3,653 |
| DeepSeek-V4-Pro | Doubao-lite | |||||
|---|---|---|---|---|---|---|
| Seed | Rubric-RL | ProRubric | Rubric-RL | ProRubric | ||
| 42 | 24.3 | 39.8 | +15.5 | 69.9 | 78.0 | +8.1 |
| 43 | 28.4 | 36.2 | +7.7 | 71.3 | 74.9 | +3.6 |
| 44 | 25.1 | 34.5 | +9.4 | 68.3 | 74.9 | +6.6 |
| Mean | 26.0 | 36.8 | +10.8 | 69.8 | 75.9 | +6.1 |
| Edit to the answer | Rubric-RL | raw-AND | Graded | ProRubric | Rubric-RL | ProRubric | Implicit | Writing |
|---|---|---|---|---|---|---|---|---|
| + appr. | + appr. | (Rubric-RL) | ||||||
| Name an item | ||||||||
| [ , ] | [ , ] | [ , ] | [ , ] | [ , ] | [ , ] | [ , ] | n.s. | |
| Name the topic | ||||||||
| [ , ] | ||||||||
| Satisfy a criterion in part | – |
| Medicine | Dialogue | Science | Writing | |||||
|---|---|---|---|---|---|---|---|---|
| Method | Score | [95% CI] | Score | [95% CI] | Score | [95% CI] | Score | [95% CI] |
| Base | 50.0 | [ , ] | 45.6 | [ , ] | 52.2 | [ , ] | 38.9 | [ , ] |
| Rubric-RL | 18.9 | — | 38.2 | — | 41.0 | — | 58.4 | — |
| ProRubric | 39.2 | [ , ] | 44.9 | [ , ] | 51.0 | [ , ] | 64.1 | [ , ] |
| raw-AND | 29.9 | [ , ] | 45.9 | [ , ] | 43.6 | [ , ] | 59.9 | [ , ] |
| Implicit, seed 42 | – | – | 43.8 | [ , ] | 44.2 | [ , ] | 66.0 | [ , ] |
| (a) Appropriateness change from the untrained model | ||||||
| Domain | Untrained | Rubric-RL | ProRubric | ProRubric Rubric-RL | ||
| score | 42 / 43 / 44 | mean | 42 / 43 / 44 | mean | mean (lowest seed) | |
| Medicine | 50.0 | / / | 2.3 | / / | 4.5 | ( ) |
| Science | 52.2 | / / | 0.4 | / / | 2.7 | ( ) |
| Dialogue | 45.6 | / / | 2.8 | / / | 0.5 | ( ) |
| Writing | 38.9 | / / | 3.1 | / / | 1.4 | ( ) |