cs.CLMar 24, 2026

Unbiased Reward Modeling from Implicit Feedback for LLM Alignment

Authors: Hao Wang, Haocheng Yang, Licheng Pan, Lei Shen, Xiaoxi Li, Yinuo Wang, Zhichao Chen, Yuan Lu, +2 more

Organizations: Peking University · Xiaohongshu Inc. · National University of Singapore · Zhejiang University

Abstract

Despite the success of reinforcement learning from human feedback (RLHF), existing reward modeling methods largely rely on explicit feedback, which is costly to collect and difficult to scale. This work studies implicit reward modeling, learning reward models from implicit user feedback, such as clicks, copies and skips. While scalable and cost-effective, implicit feedback poses two key challenges: It lacks definitive negative samples, which makes standard positive-negative classification methods inapplicable; It suffers from selection bias, where responses have heterogeneous propensities to elicit feedback, which further obscures definitive negative samples. To address these challenges, we propose ImplicitRM, which learns unbiased reward models from implicit feedback. It stratifies training samples into four latent groups using a stratification model and derives a likelihood-maximization objective that is theoretically unbiased, thereby addressing both challenges. Experiments across diverse LLM backbones and benchmark datasets validate that ImplicitRM learns accurate reward models from implicit feedback and improves performance on downstream RLHF tasks.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Your Mouse and Eyes Secretly Leak Your Preference: LLM Alignment using Implicit Feedback from Users

    Jun 18, 2026Haw-Shiuan Chang, Jeffrey Gomez, Mehul Patwari +2Large Language Model AlignmentLarge Language Model Evaluation

  2. Uncertainty-Aware Reward Modeling for Stable RLHF

    Jun 18, 2026Licheng Pan, Haocheng Yang, Haoxuan Li +7Reinforcement Learning From Human FeedbackLarge Language Model Alignment