cs.CLSep 27, 2026

RewardExplainer: Learning Reward Model Explanations from Counterfactual Preference Feedback

Authors: Jingyi He, Nier Wu, Shuang Liu, Xin Wang, Mengnan Du, Xia Hu

Organizations: The Chinese University of Hong Kong, Shenzhen · Shanghai Jiao Tong University · Carnegie Mellon University · Jilin University · Shanghai Artificial Intelligence Laboratory

Abstract

Reward models (RMs) are a key component of large language model post-training, providing reward signals for subsequent reinforcement learning. However, conventional discriminative RMs typically output only scalar scores, making it difficult to identify the response behaviors associated with their scoring decisions. Existing interpretation methods often rely on predefined high-level attributes and require repeated counterfactual interventions for each response pair to validate candidate explanations, lacking a closed-loop mechanism that uses RMs' feedback to train a reusable explainer. To address this, we propose RewardExplainer, a framework that obtains feedback from the target reward model through counterfactual rewriting and uses this feedback to further optimize the explainer. RewardExplainer generates open-ended, atomic, and intervenable natural-language scoring mechanisms, making explanations more concrete, readable, and actionable. It further converts counterfactual feedback into preference supervision, enabling the explainer to more faithfully capture the target RM's scoring preferences and sensitive behaviors than single-pass generation. Extensive experiments across multiple target RMs and explainer backbones show consistent improvements. Beyond interpretation, we use the generated mechanisms to identify potential bias patterns and construct targeted debiasing data for fine-tuning the reward model, improving robustness on reward-hacking benchmarks.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations

    Apr 7, 2025Pedro Ferreira, Wilker Aziz, Ivan TitovExplainable AI MethodsChain-of-Thought Reasoning

  2. Debiasing Reward Models via Causally Motivated Inference-Time Intervention

    Apr 30, 2026Kazutoshi Shinoda, Kosuke Nishida, Kyosuke NishidaLarge Language Model BiasLarge Language Model Alignment