Rubric Rewards from Item Response Theory
Organizations: Department of Electrical and Computer Engineering, University of British Columbia · Microsoft · School of Biomedical Engineering, University of British Columbia
Abstract
Many language tasks have no single answer that can be checked automatically. Rubrics provide criteria for judging responses to these tasks. For reinforcement learning, the resulting verdicts must be combined into a scalar reward. A common approach sums the points assigned to satisfied criteria. Distinct verdict patterns can thus receive the same reward, and the fixed points encode how much each criterion should count, not how strongly its verdict distinguishes the current rollouts. Beyond this aggregation problem, judging the full rubric needs more judge requests as the criterion count grows. To address these limitations, Rubric Response Theory (RRT) measures quality and selects criteria when rubric criteria are monotone indicators of a shared target. Rather than adding assigned points, RRT uses a two parameter item response model that treats the verdict pattern as evidence about scalar quality specific to the rubric. Under this model, its likelihood score maximizes the local signal-to-noise ratio for quality. Its Response Parameter Network (RPN) reads the prompt and criterion text to predict criterion difficulty and discrimination. As the policy distribution changes during training, RRT uses online expectation maximization to update the RPN from current rollout verdicts. With Qwen3.5-4B as the policy, RRT's macro criterion score across Medical, Science, Rubrics as Rewards Science, and RubricBench is 1.7 points above that of group relative policy optimization (GRPO). On hard and very hard criteria in Medical and Science, RRT gains 2.8 to 5.6 points over GRPO. At half the criterion budget, adaptive Fisher selection with a frozen RPN keeps the macro criterion score across four datasets within 0.1 points of GRPO with full judging. These results show RRT can reduce judge requests while remaining competitive with GRPO.
Figures & tables
| Prediction method | Science | RaR Science | Medical | RubricBench | Macro mean |
|---|---|---|---|---|---|
| Rollout evidence: ROC-AUC within each criterion | |||||
| RPN without rollout evidence | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 |
| Mean of other verdicts | 69.6 | 67.2 | 63.8 | 62.5 | 65.8 |
| RRT, | 69.7 | 67.0 | 63.9 | 62.8 | 65.8 |
| RRT | 70.0 (20.0) | 67.4 (17.4) | 64.0 (14.0) | 61.8 (11.8) | 65.8 (15.8) |
| Joint rollout and criterion ranking: pooled ROC-AUC | |||||
| Medical | Science | RaR Science | RubricBench | Macro mean | |||||
| Condition | Criterion score | Normalized points score | Criterion score | Normalized points score | Criterion score | Normalized points score | Criterion score | Criterion score | Normalized points score |
| Base policy | 64.1 | 63.3 | 76.5 | 70.9 | 68.7 | 71.0 | 74.7 | 71.0 | 70.0 |
| Vanilla GRPO | 67.7 (3.6) | 66.8 (3.5) | 78.3 (1.8) | 77.8 (6.9) | 70.6 (1.9) | 73.7 (2.7) | 77.0 (2.3) | 73.4 (2.4) | 73.8 (3.8) |
| POW3R | 70.6 (6.5) | 69.6 (6.3) | 79.0 (2.5) | 78.7 (7.8) | 70.1 (1.4) | 73.2 (2.2) | 77.4 (2.7) | 74.3 (3.3) | 74.7 (4.7) |
| DIVA | 64.9 (0.8) | 64.4 (1.1) | 78.8 (2.3) | 78.3 (7.4) | 70.6 (1.9) | 73.8 (2.8) | 75.8 (1.1) | 72.5 (1.5) | 73.1 (3.1) |
| Same RRT reward and E-step across criterion parameter settings | |||||||||
| Selection method | Medical | Science | RaR Science | RubricBench | Macro mean |
|---|---|---|---|---|---|
| Random | 10.7% | 15.2% | 5.1% | 8.9% | 10.0% |
| Discrimination ( ) | 20.0% (9.3) | 18.0% (2.7) | 5.1% (0.0) | 8.9% (0.0) | 13.0% (3.0) |
| Static Fisher | 20.0% (9.3) | 20.5% (5.3) | 18.3% (13.2) | 18.6% (9.6) | 19.3% (9.3) |
| Adaptive Fisher | 20.8% (10.1) | 22.9% (7.6) | 18.3% (13.2) | 21.9% (13.0) | 21.0% (11.0) |
| Medical | Science | RaR Science | RubricBench | Macro mean | |||||
| Condition | Criterion score | Normalized points score | Criterion score | Normalized points score | Criterion score | Normalized points score | Criterion score | Criterion score | Normalized points score |
| Vanilla GRPO | 67.7 | 66.8 | 78.3 | 77.8 | 70.6 | 73.7 | 77.0 | 73.4 | 73.8 |
| RRT, full judging (1.00) | 70.1 (2.4) | 69.8 (3.0) | 79.5 (1.2) | 79.2 (1.4) | 70.7 (0.1) | 73.8 (0.1) | 77.9 (0.9) | 74.6 (1.2) | 75.2 (1.4) |
| Random selection (0.50) | |||||||||
| Vanilla GRPO | 65.9 (1.8) | 64.0 (2.8) | 74.7 (3.6) | 74.8 (3.0) | 67.5 (3.1) | 70.3 (3.4) | 76.3 (0.7) | 71.1 (2.3) | 71.4 (2.5) |
| RRT | 67.8 (0.1) | 67.0 (0.2) | 77.2 (1.1) | 77.0 (0.8) | 70.3 (0.3) | 72.2 (1.5) | 76.5 (0.5) | 73.0 (0.5) | 73.2 (0.7) |
Appendix figures & tables37 assets
Supplementary material from the paper’s appendix.
Appendix
| Verdict included in fit (ROC-AUC) | |||||
| RPN configuration | Medical | Science | RaR Science | RubricBench | Macro mean |
| (Gaussian CDF) | 80.6 | 84.2 | 91.7 | 91.9 | 87.1 |
| (logistic CDF) | 80.5 | 84.1 | 90.1 | 90.6 | 86.3 |
| (complementary log-log) | 80.5 | 84.3 | 92.2 | 91.8 | 87.2 |
| Number of criterion parameters | |||||
| Moran’s | ||||||
|---|---|---|---|---|---|---|
| Dataset | Difficulty | T@10 | ||||
| RubricBench | 38.0 | 97.2 | 98.8 | 21.5 | 74.0 | 45.9 |
| Medical | 47.9 | 99.2 | 99.7 | 25.9 | 77.0 | 32.9 |
| Science | 47.2 | 99.2 | 99.3 | 28.2 | 85.8 | 38.2 |
| RaR Science | 38.5 | 99.1 | 99.1 | 19.4 | 74.6 | 44.1 |
| Dataset | separated | separated | Count order reversals |
|---|---|---|---|
| Medical | 83.3% | 0.0% | 0 |
| Science | 39.3% | 0.0% | 0 |
| Group size | |||||||
|---|---|---|---|---|---|---|---|
| Quantity | Reward | 2 | 4 | 8 | 16 | 24 | 32 |
| Medical | |||||||
| Tied pair share | Rubric points | 5.6 | 5.9 | 5.8 | 5.9 | 5.9 | 5.9 |
| RRT | 3.4 | 3.7 | 3.6 | 3.6 | 3.6 | 3.6 | |
| Variance share within prompts | Rubric points | 9.2 | 13.2 | 15.7 | 16.6 | 16.9 | 17.3 |
| RRT | 10.3 | 14.8 | 17.5 | 18.6 | 18.9 | 19.3 | |
| Dataset | Qwen3.5-4B | Llama-3.1-8B-Instruct | Qwen3.5-2B |
|---|---|---|---|
| Medical | 60.5 | 60.4 | 53.4 |
| Science | 69.3 | 54.6 | 52.3 |
| RaR Science | 80.5 | 52.5 | 60.6 |
| RubricBench | 73.4 | 57.9 | 56.1 |
| Medical | Science | Macro mean | ||||
|---|---|---|---|---|---|---|
| Target | Frozen | Online | Frozen | Online | Frozen | Online |
| Checkpoint aggregate | 53.2 | 55.2 (2.0) | 56.1 | 56.7 (0.6) | 54.7 | 56.0 (1.3) |
| Next policy step | 56.9 | 59.6 (2.7) | 58.9 | 59.6 (0.7) | 57.9 | 59.6 (1.7) |
| Method | Criterion loss | ROC-AUC | |
|---|---|---|---|
| Hard EM | 0 | 0.369 | 91.0 |
| Hard EM | 0.05 | 0.364 | 91.9 |
| Soft EM | 0 | 0.380 | 90.4 |
| Soft EM | 0.05 | 0.381 | 90.6 |
| Dataset | Pairwise agreement | Unanimous triples | AC1 | Uncertain agreement | |
|---|---|---|---|---|---|
| Medical | 94.6 | 91.9 | 88.3 | 90.0 | 90.5 |
| Science | 95.0 | 92.4 | 86.2 | 92.1 | 90.3 |
| RaR Science | 95.2 | 92.9 | 86.9 | 92.5 | 86.0 |
| Macro mean | 94.9 | 92.4 | 87.1 | 91.6 | 88.9 |
| Dataset | Policy | Disagreement | Unanimous triples | PRESENT | NOT_PRESENT |
|---|---|---|---|---|---|
| Medical | Qwen3.5-4B | 1.58 | 95.27 | 1.52 | 1.63 |
| Qwen3.5-2B | 1.33 | 96.01 | 0.94 | 2.15 | |
| Llama-3.1-8B-Instruct | 1.16 | 96.53 | 0.76 | 2.35 | |
| Science | Qwen3.5-4B | 1.39 | 95.83 | 2.57 | 0.95 |
| Qwen3.5-2B | 1.51 | 95.46 | 1.69 | 1.37 | |
| Llama-3.1-8B-Instruct | 1.21 | 96.36 | 0.90 | 1.89 |
| Tied pair share | Order preservation rate | Stable nonzero ordering share | ||||
|---|---|---|---|---|---|---|
| Dataset | RRT + marginal calibration | Rubric points | RRT + marginal calibration | Rubric points | RRT + marginal calibration | Rubric points |
| Medical | 1.7 | 2.5 | 82.2 | 77.5 | 80.8 | 75.6 |
| Science | 12.5 | 13.3 | 71.8 | 70.6 | 62.9 | 61.2 |
| RaR Science | 46.8 | 60.2 | 65.2 | 63.8 | 34.7 | 25.4 |
| Macro mean | 20.3 | 25.3 | 73.1 | 70.6 | 59.5 | 54.1 |
| Stable nonzero ordering share | Order flip rate | |||||
|---|---|---|---|---|---|---|
| Corruption channel | Level | Mean | Ahead | Behind | Mean | Lower |
| RRT + marginal calibration, | ||||||
| Triples with split verdicts | 0.025 | 2 | 2 | 3 | ||
| 0.050 | 2 | 1 | 3 | |||
| 0.100 | 3 | 0 | 7 | |||
| All triples | 0.025 | 2 | 3 | 2 | ||
| Cells | ||||
|---|---|---|---|---|
| Corruption channel | Structure | Mean (points) | Ahead | Behind |
| Symmetric, all criteria | None | 0 | 17 | |
| Lenient, all criteria | None | 15 | 0 | |
| Lenient, longer responses | Across rollouts | 0 | 19 | |
| Symmetric and lenient mixture | Both | 5 | 1 | |
| Symmetric, 30% of criteria | Across criteria | 19 | 0 | |
| Dataset | Corruption | Criterion score | Rubric points | POW3R | DIVA | RRT + marginal calibration | |
|---|---|---|---|---|---|---|---|
| Medical | 0.00 | 100.0 | 97.8 | 97.4 | 94.9 | 95.6 | |
| 0.10 | 68.9 | 66.6 | 64.1 | 60.4 | 72.1 | ||
| 0.20 | 49.7 | 47.7 | 44.1 | 38.8 | 55.7 | ||
| 0.30 | 30.6 | 29.3 | 26.7 | 22.7 | 35.0 | ||
| Science | 0.00 | 100.0 | 97.3 | 97.1 | 94.3 | 96.4 | |
| 0.10 | 62.2 | 60.3 | 57.8 | 52.9 | 66.5 |
| Quantity | Range across datasets |
|---|---|
| Equal pass count in two blocks of eight | 61.4% to 77.1% |
| Uncertain criteria with a majority disagreement | 25.5% to 31.7% |
| Reliability at eight rollouts, all criteria | 93.6% to 96.0% |
| Reliability at eight rollouts, uncertain criteria | 70.9% to 71.9% |
| Difficulty correlation after attenuation adjustment | 38.5% to 52.4% |
| Dataset | Rubric points | RRT | Difference |
|---|---|---|---|
| Medical | 26.6 | 27.1 | |
| Science | 21.5 | 25.1 | |
| RaR Science | 12.4 | 27.5 | |
| RubricBench | 25.1 | 27.5 |
| Rank correlation quintile | RRT gain |
|---|---|
| 1, lowest rank correlation | |
| 2 | |
| 3 | |
| 4 | |
| 5, highest rank correlation |
| Pairwise dependence | Residual | Residual share | Redundant share by | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Dataset | Explained | Observed | Bootstrap | Redundant | Deficient | ||||
| Medical | 4.3 | 6.3 | 72.8 | 1.7 | 1.8 | 6.0 | 6.0 | 3.9 (4.2) | 23.5 (11.6) |
| Science | 9.0 | 11.2 | 76.4 | 3.4 | 3.8 | 6.8 | 7.7 | 4.2 (3.6) | 20.2 (12.2) |
| RaR Science | 12.1 | 18.1 | 75.4 | 6.8 | 10.8 | 5.6 | 11.2 | 3.1 (2.7) | 13.7 (10.3) |
| RubricBench | 10.6 | 14.6 | 85.3 | 6.2 | 7.8 | 5.1 | 7.5 | 3.4 (4.1) | 16.4 (10.7) |
| Macro mean | 9.0 | 12.5 | 77.5 | 4.5 | 6.1 | 5.9 | 8.1 | 3.7 (3.7) | 18.5 (11.2) |
| Condition | Medical | Science | RaR Science | RubricBench | Macro mean |
|---|---|---|---|---|---|
| Qwen3.5-2B | |||||
| Base policy | 37.2 | 44.4 | 58.8 | 57.7 | 49.5 |
| Vanilla GRPO | 53.8 (16.7) | 59.9 (15.4) | 63.1 (4.3) | 60.6 (2.9) | 59.4 (9.8) |
| RRT | 53.3 (16.1) | 59.1 (14.7) | 63.7 (4.9) | 60.8 (3.1) | 59.2 (9.7) |
| Llama-3.1-8B-Instruct | |||||
| Base policy | 31.1 | 29.8 | 46.7 | 61.9 | 42.4 |
| Condition | HealthBench | ResearchQA | Macro mean |
| Qwen3.5-4B | |||
| Base policy | 58.8 | 78.9 | 68.9 |
| Vanilla GRPO | 59.8 (1.0) | 79.7 (0.8) | 69.8 (0.9) |
| RRT | 60.9 (2.1) | 80.0 (1.1) | 70.5 (1.6) |
| Qwen3.5-2B | |||
| Base policy | 42.7 | 64.5 | 53.6 |
| Selection method | Medical | Science | RaR Science | RubricBench | Macro mean |
|---|---|---|---|---|---|
| Random | 10.7% | 11.5% | 4.9% | 8.8% | 9.0% |
| Discrimination ( ) | 16.8% (6.2) | 14.2% (2.7) | 4.9% (0.0) | 8.8% (0.0) | 11.2% (2.2) |
| Static Fisher | 19.3% (8.7) | 20.7% (9.3) | 18.2% (13.3) | 19.9% (11.1) | 19.6% (10.6) |
| Adaptive Fisher | 19.3% (8.7) | 21.7% (10.2) | 18.2% (13.3) | 14.7% (5.9) | 18.5% (9.5) |
| Dataset | Source pool | Training | Validation | Test | Split role |
|---|---|---|---|---|---|
| Medical | 29,681 | 5,000 | 500 | 500 | Validation selects checkpoints. Test gives final results. |
| Science | 29,418 | 5,000 | 500 | 500 | Validation selects checkpoints. Test gives final results. |
| RaR Science | 22,917 | 5,000 | 500 | 500 | Validation selects checkpoints. Test gives final results. |
| RubricBench | 1,147 | 847 | 150 | 150 | Validation selects checkpoints. Test gives final results. |
| HealthBench | 5,000 | 0 | 0 | 5,000 | Full set for Medical generalization. |
| ResearchQA | 21,414 | 0 | 0 | 21,414 | Full set for Science generalization. |
| Experiment case | Training role | Evaluation role |
|---|---|---|
| RPN experiments | Training prompts fit the RPNs used in policy training. | Validation selects those checkpoints. |
| Policy experiments | Training prompts update the policy. | Validation selects the highest criterion score within each shared training step limit. Test gives final policy results. |
| Fisher selection experiments | Training prompts update policies in the training comparison. | Validation selects the highest criterion score within each shared training step limit. Test gives final policy results for the training comparison. |
| Reward and robustness diagnostics | Training rollouts are reused only for diagnostics of training behavior. | Other diagnostics use the test set. They do not select checkpoints or give final policy results. |
| Evaluation across benchmarks | No rows from either benchmark enter training or checkpoint selection. | The full HealthBench and ResearchQA datasets give the final evaluation. |
| Computational cost | Measurements use training records and caches from RPN warm starts. | Checkpoint evaluation timings enter where reported. Test responses are not used. |
| Setting | Value |
|---|---|
| Job layout | Eight nodes for each full batch, with one independent run per node. Each node has eight NVIDIA A100 80 GB GPUs. |
| Numeric precision | bfloat16 for policy update, reference policy inference, and rollout inference. |
| Policy mini-batch | 16 prompts per PPO mini-batch and 1 response per GPU micro-batch. |
| Sequence lengths | Maximum prompt length 6,144 tokens and maximum response length tokens. |
| Sampling | Temperature 1.0, top- 1.0, top- disabled, repetition penalty 1.1 for Qwen and 1.0 for Llama. |
| GRPO loss | Clipped surrogate with and mean token loss within each sequence averaged across sequences. |
| Policy | Condition | Medical | Science | RaR Science | RubricBench |
|---|---|---|---|---|---|
| Primary comparison | |||||
| Qwen3.5-4B | Vanilla GRPO | 120 | 130 | 105 | 50 |
| RRT + frozen RPN | 145 | 150 | 120 | 45 | |
| RRT + online RPN | 115 | 110 | 110 | 50 | |
| Qwen3.5-2B | Vanilla GRPO | 100 | 150 | 120 | 45 |
| RRT | 125 | 150 | 110 | 45 | |
| Medical | Science | |||
| Method | Judge requests | Input tokens | Judge requests | Input tokens |
| Vanilla GRPO | 7,716 | 11.1 | 6,908 | 16.3 |
| RRT, full judging (1.00) | 7,685 | 11.1 | 6,895 | 16.5 |
| RRT, adaptive Fisher | ||||
| 0.95 | 7,447 | 10.8 | 6,670 | 16.1 |
| 0.80 | 6,265 | 9.1 | 5,622 | 12.8 |
| Quantity | Medical | Science |
|---|---|---|
| Completed judge requests | 6,728 | 4,984 |
| Input tokens per request | ||
| mean | 1,137 | 3,800 |
| median | 984 | 1,216 |
| p10 to p90 | 496 to 1,664 | 640 to 2,846 |
| maximum | 29,399 | 32,115 |
| Method | Added seconds | Overhead of the policy step |
|---|---|---|
| Vanilla GRPO | 0 | 0% |
| RRT + | ||
| batch pass rate | 0.10 | 0.004% |
| cached pass rate | 0.10 | 0.004% |
| frozen RPN | 0.17 | 0.007% |
| online RPN | 2.79 | 0.123% |
| Seconds per policy step, median with p10 to p90 | ||||||
| Condition | Generation | Judging | Log probability | Policy update | Total policy step | p10 to p90 of policy step |
| Medical | ||||||
| Vanilla GRPO | 460 | 1,451 | 36 | 107 | 2,274 | 1,136 to 3,603 |
| RRT + | ||||||
| frozen RPN | 647 | 1,457 | 38 | 124 | 2,336 | 1,815 to 3,092 |
| online RPN | 703 | 1,453 | 66 | 217 | 2,454 | 1,628 to 4,193 |
| Dataset | Optimizer steps | GPU hours, 0.6B embedder | GPU hours, 8B embedder |
|---|---|---|---|
| Medical | 3,000 | 5.30 | 7.64 |
| Science | 1,254 | 2.48 | 4.71 |
| RaR Science | 1,254 | 1.12 | 2.33 |
| RubricBench | 1,067 | 0.89 | 1.57 |
| Policy | Added deployment parameters | Added inference components |
|---|---|---|
| Base policy | 0 | None |
| Policy trained with RRT | 0 | None |