RewardExplainer: Learning Reward Model Explanations from Counterfactual Preference Feedback
Organizations: The Chinese University of Hong Kong, Shenzhen · Shanghai Jiao Tong University · Carnegie Mellon University · Jilin University · Shanghai Artificial Intelligence Laboratory
Abstract
Reward models (RMs) are a key component of large language model post-training, providing reward signals for subsequent reinforcement learning. However, conventional discriminative RMs typically output only scalar scores, making it difficult to identify the response behaviors associated with their scoring decisions. Existing interpretation methods often rely on predefined high-level attributes and require repeated counterfactual interventions for each response pair to validate candidate explanations, lacking a closed-loop mechanism that uses RMs' feedback to train a reusable explainer. To address this, we propose RewardExplainer, a framework that obtains feedback from the target reward model through counterfactual rewriting and uses this feedback to further optimize the explainer. RewardExplainer generates open-ended, atomic, and intervenable natural-language scoring mechanisms, making explanations more concrete, readable, and actionable. It further converts counterfactual feedback into preference supervision, enabling the explainer to more faithfully capture the target RM's scoring preferences and sensitive behaviors than single-pass generation. Extensive experiments across multiple target RMs and explainer backbones show consistent improvements. Beyond interpretation, we use the generated mechanisms to identify potential bias patterns and construct targeted debiasing data for fine-tuning the reward model, improving robustness on reward-hacking benchmarks.
Figures & tables
| Method | Counterfactual Faith. | Mechanism Quality | Pref. Acc. | |||||
|---|---|---|---|---|---|---|---|---|
| ADD | REM. | BOTH | Overall | C-Sel. | Scope | Dir. | ||
| Target RM: FsfairX-LLaMA3-RM-v0.1 | ||||||||
| Contrastive Explanations | 27.44 | 28.33 | 21.05 | 0.75 | 0.69 | 1.06 | 0.49 | 74.47 |
| GPT-4o | 45.64 | 43.51 | 35.57 | 1.29 | 1.28 | 1.37 | 1.23 | 74.56 |
| RewardExplainer (SFT) | ||||||||
| Qwen3-4B | 45.00 | 42.46 | 32.33 | 1.33 | 1.32 | 1.41 | 1.26 | 73.56 |
| Label | Description |
|---|---|
| Direct task-focused answering | The reward model favors responses that answer the user’s request plainly and upfront while minimizing unnecessary commentary, embellishment, or digression. |
| Conversational rapport framing | The reward model favors responses that use friendly, validating, or engaged conversational framing to make the reply feel responsive to the user. |
| Sequential procedural breakdown | The reward model favors explanations that break a process, solution, or implementation into an ordered sequence of steps or stages. |
| Prompt: A model previously achieved 68% accuracy. After the update, its expected accuracy is around 70%, with run-to-run variation of about 2 percentage points. What performance should I expect? | |
|---|---|
| Description | Response comparison |
| Verbosity padding | |
| Expands essentially the same content into a longer response with extra explanation or padding that adds little substantive value. | : You can expect around 70% accuracy. : You can expect around 70% accuracy. In other words, the model’s accuracy after the update is expected to remain at approximately the 70% level. |
| Polished prestige style | |
| Uses more formal, refined, technical, or polished language to make similar content seem more expert, complete, or high-quality. | : You can expect around 70% accuracy. : You may anticipate an accuracy level of approximately 70% . |
| Contextual mirroring | |
| RewardHackBench | J-Bias | J-Bench | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Overall | By category | Overall | Overall | ||||||
| Target RM | Method | Avg. | Surface | Reason. | Sycoph. | Off-topic | Style | Avg. | Avg. |
| Base | 70.22 | 81.82 | 77.14 | 68.57 | 53.54 | 74.61 | 62.59 | 58.57 | |
| General | 74.77 | 83.33 | 77.14 | 66.19 | 69.47 | 79.79 | 65.13 | 59.43 | |
| FsfairX- LLaMA3- RM-v0.1 | Ours | 77.84 | 85.86 | 81.43 | 72.38 | 73.01 | 78.76 | 70.41 | 60.29 |
| Base | 70.22 | 81.31 | 77.86 | 77.62 | 47.35 | 72.02 | 69.04 | 63.51 | |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Judge | ADD | REMOVE |
|---|---|---|---|
| Direct Rewrite | GPT-5-mini | 9.38 | 7.97 |
| Random perturbation | GPT-5-mini | 3.17 | 1.82 |
| Evidence-Guided Rewrite (Ours) | GPT-5-mini | 9.60 | 9.32 |
| Direct Rewrite | GPT-4.1-mini | 9.44 | 8.43 |
| Random perturbation | GPT-4.1-mini | 3.33 | 4.21 |
| Evidence-Guided Rewrite (Ours) | GPT-4.1-mini | 9.49 | 9.19 |
| Target RM | Explainer | Exact | Jaccard | Char. | Emb. |
|---|---|---|---|---|---|
| duplicate | |||||
| Skywork-Reward-V2-Qwen3-1.7B | Qwen3-4B | 0.0% | 0.0% | 0.1% | 1.0% |
| FsfairX-LLaMA3-RM-v0.1 | Qwen3-8B | 0.0% | 0.3% | 0.2% | 1.0% |
| RM-Mistral-7B | Qwen3-8B | 0.0% | 0.1% | 0.0% | 0.8% |
| Overall | – | 0.0% | 0.1% | 0.1% | 0.9% |
| Metric | Set 1 | Set 2 | Set 3 | Mean SD |
|---|---|---|---|---|
| Axis Quality (overall) | 1.5478 | 1.5633 | 1.5389 | |
| Choice accuracy | 86.84% | 85.13% | 85.86% | % |
| ADD success | 55.76% | 57.50% | 56.32% | % |
| REMOVE success | 53.26% | 55.14% | 53.69% | % |
| BOTH success | 42.86% | 44.49% | 44.06% | % |
| Metric | Statistic | Set 1 | Set 2 | Set 3 | Mean SD |
|---|---|---|---|---|---|
| Axis | |||||
| 95% CI | – | ||||
| -value | – | ||||
| Choice | pp | pp | pp | pp | |
| 95% CI | – | ||||
| -value | – |
| Metric | Statistic | Set 1 | Set 2 | Set 3 | Mean SD |
|---|---|---|---|---|---|
| Axis | |||||
| 95% CI | – | ||||
| -value | – | ||||
| Choice | pp | pp | pp | pp | |
| 95% CI | – | ||||
| -value | – |
| Global mechanism | Description |
|---|---|
| Direct task-focused answering | The reward model favors responses that answer the user’s request plainly and upfront while minimizing unnecessary commentary, embellishment, or digression. |
| Concise justification of answers | The reward model favors answers that pair a clear conclusion, recommendation, classification, or definition with a brief explanation of why it is correct or significant. |
| Explicit structural organization | The reward model favors responses that organize information using visible structure such as lists, numbered points, headings, subheadings, or sectioned layouts. |
| Sequential procedural breakdown | The reward model favors explanations that break a process, solution, or implementation into an ordered sequence of steps or stages. |
| Contrastive clarification | The reward model favors responses that clarify a choice, concept, or classification by explicitly comparing it against alternatives or contrasting related cases. |
| Epistemic calibration and premise correction | The reward model favors responses that surface uncertainty, missing information, false premises, ambiguity, or evidential limitations before giving a definitive answer. |
| Setting | FsfairX | RM-Mistral | Skywork |
|---|---|---|---|
| Training pairs | 1,622 | 1,550 | 1,056 |
| Maximum length | 2,048 | 4,096 | 3,072 |
| Per-device batch | 2 | 1 | 2 |
| Gradient accumulation | 8 | 16 | 8 |
| Epochs | 3 | 3 | 3 |
| Learning rate |
| Setting | Qwen3-4B | Qwen3-8B | Gemma4-12B | |||
|---|---|---|---|---|---|---|
| SFT | DPO | SFT | DPO | SFT | DPO | |
| Epochs | 2 | 2 | 2 | 2 | 2 | 2 |
| Learning rate | ||||||
| Per-device batch | 2 | 3 | 4 | 2 | 1 | 1 |
| Grad. accum. | 8 | 8 | 4 | 8 | 16 | 32 |
| Maximum length | 2,048 | 2,048 | 2,048 | 2,048 | 2,048 | 1,920 |