Organizations: University of Science and Technology of China, Hefei, China · SenseTime Research, Shanghai, China · Institute of Artificial Intelligence, Hefei Comprehensive National Science Center
We identify a fundamental mismatch in empathetic reinforcement learning: support priorities evolve with the dialogue state, yet existing methods typically optimize predefined reward specifications that remain fixed across turns. To model these evolving support priorities, we organize empathetic support along cognitive, affective, and proactive empathy, and propose Context-Adaptive Rubric Evolution (CARE). At each turn, CARE generates a context-adaptive rubric by adjusting both the weights of these three empathy dimensions and their fine-grained evaluation criteria. The rubric generator is trained with turn-level rubric supervision and human preference data through supervised fine-tuning followed by preference-based reinforcement learning, and then serves as an adaptive reward interface for online empathetic RL. Integrated with both RLVER and MICA, CARE achieves state-of-the-art performance across SentientBench, EQBench3, and EMPA under three independent LLM judges. Notably, on EMPA, CARE improves EPM-Idx over the strongest baseline by at least 13 points under all three judges, including an increase from 28.11 to 83.54 under Gemini-2.5-Pro. Further analyses show that learned rubric priorities systematically vary across dialogue stages and user emotions, demonstrating that CARE adapts what is rewarded as support needs evolve.
Figures & tables
Figure 1: Evolving support priorities in empathetic reinforcement learning. (a) CARE constructs a context-adaptive rubric over three empathy dimensions— cognitive, affective, and proactive —with state-dependent weights and fine-grained criteria. (b) CARE consistently outperforms prior methods across benchmarks and judges. Notably, on EMPA ( Zhang et al., 2026b ) , our CARE improves EPM-Idx over the strongest baseline with double-digit gains under all three LLM judges.
Figure 2: Overview of CARE . Turn-level supervision and human response preferences train a context-conditioned rubric generator. The frozen generator then supplies adaptive evaluation criteria for online dialogue RL, with one rubric shared by candidate responses from the same context.
SentientBench
EQBench3
EMPA
Model
Score ↑
Succ. (%) ↑
Fail (%) ↓
Overall ↑
EPM-Idx ↑
Judge: DeepSeek-V4-Pro
Qwen2.5-7B-Instruct
9.14 ± 1.48
3.33 ± 0.94
87.67 ± 0.47
42.53 ± 2.18
15.38 ± 1.06
PERM ( Wang et al., 2026b )
9.57 ± 0.62
1.67 ± 0.94
86.67 ± 1.25
49.07 ± 1.03
17.21 ± 1.17
KARDIA-R1 ( Yuan et al., 2026 )
25.45 ± 2.46
6.67 ± 1.70
58.33 ± 2.87
28.75 ± 1.21
11.42 ± 0.34
RLVER ( Wang et al., 2026c )
46.53 ± 2.13
22.00 ± 0.00
30.33 ± 4.50
44.23 ± 2.23
26.05 ± 3.37
Table 1: Main comparison across SentientBench, EQBench3, and EMPA under three LLM judges. Additional results for EQBench3 and EMPA are provided in § F . Standard deviations are computed over three random seeds; bold and underline mark the best and second-best results.
Model
Overall
DoI
ER
DE
WRM
HL
PEI
SD
GPT-5.5
83.35
17.57
16.98
16.64
14.32
16.72
16.72
16.16
Gemini-2.5-Pro
82.65
17.71
16.73
17.00
15.31
17.31
15.81
15.85
Claude-Sonnet-4.6
79.30
17.52
16.48
15.38
13.35
15.54
14.00
13.88
DeepSeek-V4-Pro-1.6T
77.75
16.91
16.14
15.69
13.73
16.31
13.96
13.69
Qwen3.7-Max
76.60
17.00
15.51
15.65
13.23
14.92
14.58
13.77
GLM-5.2-753B
74.90
16.84
15.43
14.62
12.46
15.12
13.35
12.92
Table 2: Scaling and frontier comparison on EQBench3 under DeepSeek-V4-Pro. We compare Base and CARE -enhanced Qwen3 models from 8B to 32B, together with strong frontier LLMs.
Table 5Figure 6
Figure 5: Support-strategy occurrence rates.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Cognitive-dominant rubric case.
Figure 7: Affective-dominant rubric case.
Figure 8: Proactive-dominant rubric case.
Figure 9: Strict-maximum empathy-dimension proportions by dialogue turn for SFT (left) and preference RL (right). Each dialogue-turn first aggregates repeated rubric rollouts; curves then report dimension-level proportions across dialogues.
Figure 10: Emotion-conditioned tactic use per method. Each cell reports the fraction of scenario-turns associated with a given emotion label that contain the corresponding MINT tactic.
Model
Overall
DoI
ER
DE
WRM
HL
PEI
SD
Judge: DeepSeek-V4-Pro
Qwen2.5-7B-Instruct
42.53 ± 2.18
9.35 ± 0.49
8.97 ± 0.47
10.09 ± 0.69
9.41 ± 0.19
8.59 ± 0.68
8.85 ± 0.98
7.56 ± 0.91
PERM ( Wang et al., 2026b )
49.07 ± 1.03
10.97 ± 0.15
10.26 ± 0.10
11.98 ± 0.32
10.63 ± 0.32
10.36 ± 0.43
10.90 ± 0.40
9.58 ± 0.43
KARDIA-R1 ( Yuan et al., 2026 )
28.75 ± 1.21
6.00 ± 0.36
6.04 ± 0.23
8.21 ± 0.12
9.61 ± 0.09
6.80 ± 0.46
6.15 ± 0.27
5.52 ± 0.23
RLVER ( Wang et al., 2026c )
44.23 ± 2.23
9.71 ± 0.62
9.39 ± 0.41
10.88 ± 0.35
10.50 ± 0.60
9.36 ± 0.43
9.23 ± 0.47
8.14 ± 0.49
MICA (rep.) ( Zhang et al., 2026a )
51.47 ± 0.58
11.51 ± 0.20
10.92 ± 0.16
11.74 ± 0.19
11.18 ± 0.09
11.22 ± 0.28
9.88 ± 0.49
9.11 ± 0.66
Appendix
Table 6: Complete EQBench3 results under three LLM judges, including Overall and all seven component scores. Values are means ± standard deviations over three runs; all metrics are higher-is-better, and bold and underline mark the best and second-best means.
Model
EPM-Idx
Outcome
Outcome
RDI
Etot
Snet
Judge: DeepSeek-V4-Pro
Qwen2.5-7B-Instruct
15.38 ± 1.06
3.25 ± 0.76
7.5 ± 1.5
0.0 ± 0.0
2.3 ± 0.9
PERM ( Wang et al., 2026b )
17.21 ± 1.17
6.39 ± 2.14
10.6 ± 1.7
0.3 ± 0.4
8.2 ± 4.8
KARDIA-R1 ( Yuan et al., 2026 )
11.42 ± 0.34
0.81 ± 0.34
2.4 ± 1.0
0.0 ± 0.0
0.0 ± 0.0
RLVER ( Wang et al., 2026c )
26.05 ± 3.37
15.34 ± 4.72
17.3 ± 4.5
5.2 ± 2.9
23.5 ± 7.3
Appendix
Table 7: Complete EMPA EPM-Idx and outcome results under three LLM judges.
Model
Efficiency
Stability
Efficiency
Rho
Sproj
Tau
Stability
Rpos
Align
Pen
Judge: DeepSeek-V4-Pro
Qwen2.5-7B-Instruct
25.67 ± 1.15
0.0 ± 0.0
0.0 ± 0.0
77.0 ± 3.4
22.35 ± 2.38
17.7 ± 1.5
22.3 ± 2.4
27.2 ± 3.3
PERM ( Wang et al., 2026b )
23.19 ± 1.07
0.2 ± 0.2
0.2 ± 0.2
69.3 ± 3.3
25.04 ± 2.25
21.3 ± 2.2
25.2 ± 2.3
28.7 ± 2.4
KARDIA-R1 ( Yuan et al., 2026 )
30.91 ± 0.50
0.0 ± 0.0
0.0 ± 0.0
92.7 ± 1.5
12.29 ± 0.76
8.1 ± 0.9
12.9 ± 0.7
15.9 ± 0.7
RLVER ( Wang et al., 2026c )
17.86 ± 3.15
2.0 ± 1.2
2.0 ± 1.1
49.6 ± 8.8
40.87 ± 4.36
37.2 ± 4.5
39.4 ± 4.2
46.0 ± 5.1
Appendix
Table 8: Complete EMPA results under three LLM judges (continued).
Model
Selections
Preference (%) ↑
MICA (rep.) ( Zhang et al., 2026a )
104
23.9
RLVER ( Wang et al., 2026c )
107
24.6
CARE(R)
224
51.5
Appendix
Table 9: Blinded human preference evaluation. fifteen annotators each evaluate 30 randomly assigned contexts. Model identities are hidden and response order is randomly shuffled.
Large language models have demonstrated conversational capabilities, yet empathetic competence remains challenging. Empathetic support is inherently multi-turn and path-dependent: users disclose concerns gradually, emotions evolve over time, and early responses shape trust and receptivity. Reinforcement learning with verifiable emotion rewards provides scalable supervision for long-horizon interactions. However, existing methods evolve the dialogue policy while keeping its training interaction distribution fixed, creating a mismatch between policy competence and training experience. We introduce a dual-loop self-evolution framework driven by verifiable emotion feedback. With the user simulator and verifier frozen, the inner loop optimizes the multi-turn policy using continuous emotion rewards, while the outer loop uses the same outcomes to estimate policy-relative interaction utility and adapt experience. To obtain estimates from sparse, stochastic rollouts, the framework holds the scenario and interaction state constant within each group and prioritizes conditions whose group pass rates lie near the policy's competence boundary. A hierarchical controller shares evidence across support intents, while uncertainty-guided exploration and uniform rehearsal prevent premature exclusion. The resulting distribution generates trajectories, closing both loops without increasing the rollout budget. On SAGE, our framework raises Qwen3-8B Overall from 53.87 to 79.24 and outperforms protocol-matched uniform emotion-reward reinforcement learning by 7.23 points.
Yi Wei, Shuo Jiang, Huaixia Dou +5
Qwen DianJin Team, Alibaba Cloud Computing · Beihang University · School of Computer Science and Technology, Soochow University
Reinforcement learning from verifiable emotion rewards RLVER has produced language models with strong empathetic performance, evaluated on benchmarks that assume cooperative, honest users. Yet real emotional interactions systematically violate this assumption: users gaslight, escalate, and pressure AI systems for unconditional validation, dynamics that cooperative benchmarks cannot surface. We construct the Adversarial Empathy Benchmark AEB and introduce the Emotional Consistency Score ECS to evaluate empathetic robustness under adversarial conditions. AEB comprises six psychologically grounded adversarial trajectory types with discriminative reward structures that penalize formulaic responses; ECS formally disentangles a model's capacity to track user emotional states from its capacity to improve them. In a controlled experiment across eight scenario-matched conditions (think and no-think conditions on 2 RLVER models, and 2 base models (Qwen 1.5B and 7B) with 480 adversarial dialogues), RLVER-PPO-Think substantially outperforms the same-scale untuned baseline (0.963 vs. 0.761, p<0.001,r=0.688), with zero dialogue collapses and 47% higher hidden-intention detection. However, ECS remains nearly flat and is not significantly different for RLVER-PPO-Think versus Base-7B-Think (p=0.650): RL training improves emotional responsiveness without measurable gains in observable state tracking. We interpret the ECS--FS (Final Score) gap as a behavioral/legibility dissociation inside this simulator family, not as evidence about internal understanding or clinical readiness.
Deeraj S K, Sadhana Devarajan, Krishna Mehra +1
Department of Artificial Intelligence, Sardar Vallabhbhai National Institute of Technology, Surat, India
As Large Language Models (LLMs) are increasingly deployed in long-term interactions with users, empathy has become an increasingly important capability. However, existing research overlooks the influence of users' personality traits on empathetic strategies during long-term interactions. To address this gap, we introduce the task of personalized empathy, which focuses on adapting empathetic strategies according to users' personalized characteristics derived from history. To study and enhance this capability, we construct PersonaEmp, a personalized empathy dataset built from long-term user-AI interactions, featuring rich user histories, persona information, and empathy-seeking queries. We further propose PereGRM, a reward modeling framework that combines the empathy evaluation structure with dynamic evaluation criteria generation for fine-grained reward modeling. Experimental results across different settings and multiple judge models show that PereGRM consistently achieves the strongest performance improvements, indicating its effectiveness for enhancing personalized empathetic capabilities.
Wuqiang Zheng, Chengbing Wang, Yilin Yang +6
University of Science and Technology of China · 2Huawei Technologies · 3China Academy of Cyber