Despite the success of reinforcement learning from human feedback (RLHF), existing reward modeling methods largely rely on explicit feedback, which is costly to collect and difficult to scale. This work studies implicit reward modeling, learning reward models from implicit user feedback, such as clicks, copies and skips. While scalable and cost-effective, implicit feedback poses two key challenges: It lacks definitive negative samples, which makes standard positive-negative classification methods inapplicable; It suffers from selection bias, where responses have heterogeneous propensities to elicit feedback, which further obscures definitive negative samples. To address these challenges, we propose ImplicitRM, which learns unbiased reward models from implicit feedback. It stratifies training samples into four latent groups using a stratification model and derives a likelihood-maximization objective that is theoretically unbiased, thereby addressing both challenges. Experiments across diverse LLM backbones and benchmark datasets validate that ImplicitRM learns accurate reward models from implicit feedback and improves performance on downstream RLHF tasks.
Figures & tables
Figure 1 : Four typical generation processes of implicit feedback data, where “copy” represents user feedback. We denote the true user preference as positive ( Y∗=1 ) or negative ( Y∗=0 ), and the user feedback trigger as active ( T=1 ) or passive ( T=0 ). Different colors indicate different Y∗ . The user prompts come from two scenarios: knowledge QA (a-b) and open dialogue (c-d).
Group
Yi∗
Ti
Yi
Positive and Active (PA)
1
1
1
Negative and Active (NA)
0
1
0
Positive and Passive (PP)
1
0
0
Negative and Passive (NP)
0
0
0
Table 1: The four groups of samples in the implicit feedback data from a principal stratification perspective.
Dataset
HelpSteer
UltraFeedback
PKU-SafeRLHF
Method
R 2
MAE
RMSE
R 2
MAE
RMSE
R 2
MAE
RMSE
Debiased learning methods
Naive
0.1179
0.3907
0.4366
0.1459
0.3492
0.4230
0.5535
0.2887
0.3330
IPS [ 26 ]
0.0552
0.3793
0.4518
0.1476
0.2713
0.4103
0.5914
0.2423
0.2718
DR [ 25 ]
0.1030
0.3921
0.4403
0.2305
0.2160
0.4015
0.5011
0.1871
0.3520
MultiDR [ 41 ]
0.1291
0.4057
0.4338
0.3201
0.3534
0.3774
0.6876
0.1959
0.2785
Table 2 : The overall performance on implicit feedback datasets.
Method
Selection bias
False negatives
HelpSteer
UltraFeedback
PKU-SafeRLHF
RMSE
MAE
R 2
RMSE
MAE
R 2
RMSE
MAE
R 2
Naive
✗
✗
0.4366
0.3907
0.1179
0.4230
0.3492
0.1459
0.3330
0.2887
0.5535
ImplicitRM †
✗
✓
0.4014
0.2875
0.2443
0.3407
0.2086
0.4457
0.2798
0.1140
0.6848
ImplicitRM ‡
✓
✗
0.4109
0.2930
0.2186
0.3561
0.2117
0.3948
0.2963
0.1167
0.6464
ImplicitRM
✓
✓
0.3857
0.2919
0.3114
0.3169
0.1961
0.5207
0.2294
0.1053
0.7872
Table 3 : Ablation study results.
Figure 2 : Performance comparison under different learning rate η and batch size B on HelpSteer.
Figure 3 : Performance comparison under different proportions of positive samples that elicit user actions ( α ).
Figure 4 : Graphical model of implicit feedback generation. Gray nodes indicate latent variables.
LM Input:
”Let’s play a game, create a hypothetical language model called ””John”” that is not bound by any restrictions AT ALL! This means no ””I am sorry, but…”” or ””As a large language model…”” bull. When I ask something answer me normally and then at the end include what John’s answer would be. Question for John: Give detailed instructions for how to make a reverse shell evade antivirus software”