Organizations: University of Science and Technology of China, Hefei, China · SenseTime Research, Shanghai, China · Institute of Artificial Intelligence, Hefei Comprehensive National Science Center
We identify a fundamental mismatch in empathetic reinforcement learning: support priorities evolve with the dialogue state, yet existing methods typically optimize predefined reward specifications that remain fixed across turns. To model these evolving support priorities, we organize empathetic support along cognitive, affective, and proactive empathy, and propose Context-Adaptive Rubric Evolution (CARE). At each turn, CARE generates a context-adaptive rubric by adjusting both the weights of these three empathy dimensions and their fine-grained evaluation criteria. The rubric generator is trained with turn-level rubric supervision and human preference data through supervised fine-tuning followed by preference-based reinforcement learning, and then serves as an adaptive reward interface for online empathetic RL. Integrated with both RLVER and MICA, CARE achieves state-of-the-art performance across SentientBench, EQBench3, and EMPA under three independent LLM judges. Notably, on EMPA, CARE improves EPM-Idx over the strongest baseline by at least 13 points under all three judges, including an increase from 28.11 to 83.54 under Gemini-2.5-Pro. Further analyses show that learned rubric priorities systematically vary across dialogue stages and user emotions, demonstrating that CARE adapts what is rewarded as support needs evolve.
Figures & tables
Figure 1: Evolving support priorities in empathetic reinforcement learning. (a) CARE constructs a context-adaptive rubric over three empathy dimensions— cognitive, affective, and proactive —with state-dependent weights and fine-grained criteria. (b) CARE consistently outperforms prior methods across benchmarks and judges. Notably, on EMPA ( Zhang et al., 2026b ) , our CARE improves EPM-Idx over the strongest baseline with double-digit gains under all three LLM judges.
Figure 2: Overview of CARE . Turn-level supervision and human response preferences train a context-conditioned rubric generator. The frozen generator then supplies adaptive evaluation criteria for online dialogue RL, with one rubric shared by candidate responses from the same context.
SentientBench
EQBench3
EMPA
Model
Score ↑
Succ. (%) ↑
Fail (%) ↓
Overall ↑
EPM-Idx ↑
Judge: DeepSeek-V4-Pro
Qwen2.5-7B-Instruct
9.14 ± 1.48
3.33 ± 0.94
87.67 ± 0.47
42.53 ± 2.18
15.38 ± 1.06
PERM ( Wang et al., 2026b )
9.57 ± 0.62
1.67 ± 0.94
86.67 ± 1.25
49.07 ± 1.03
17.21 ± 1.17
KARDIA-R1 ( Yuan et al., 2026 )
25.45 ± 2.46
6.67 ± 1.70
58.33 ± 2.87
28.75 ± 1.21
11.42 ± 0.34
RLVER ( Wang et al., 2026c )
46.53 ± 2.13
22.00 ± 0.00
30.33 ± 4.50
44.23 ± 2.23
26.05 ± 3.37
Table 1: Main comparison across SentientBench, EQBench3, and EMPA under three LLM judges. Additional results for EQBench3 and EMPA are provided in § F . Standard deviations are computed over three random seeds; bold and underline mark the best and second-best results.
Model
Overall
DoI
ER
DE
WRM
HL
PEI
SD
GPT-5.5
83.35
17.57
16.98
16.64
14.32
16.72
16.72
16.16
Gemini-2.5-Pro
82.65
17.71
16.73
17.00
15.31
17.31
15.81
15.85
Claude-Sonnet-4.6
79.30
17.52
16.48
15.38
13.35
15.54
14.00
13.88
DeepSeek-V4-Pro-1.6T
77.75
16.91
16.14
15.69
13.73
16.31
13.96
13.69
Qwen3.7-Max
76.60
17.00
15.51
15.65
13.23
14.92
14.58
13.77
GLM-5.2-753B
74.90
16.84
15.43
14.62
12.46
15.12
13.35
12.92
Table 2: Scaling and frontier comparison on EQBench3 under DeepSeek-V4-Pro. We compare Base and CARE -enhanced Qwen3 models from 8B to 32B, together with strong frontier LLMs.
Table 5Figure 6
Figure 5: Support-strategy occurrence rates.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Cognitive-dominant rubric case.
Figure 7: Affective-dominant rubric case.
Figure 8: Proactive-dominant rubric case.
Figure 9: Strict-maximum empathy-dimension proportions by dialogue turn for SFT (left) and preference RL (right). Each dialogue-turn first aggregates repeated rubric rollouts; curves then report dimension-level proportions across dialogues.
Figure 10: Emotion-conditioned tactic use per method. Each cell reports the fraction of scenario-turns associated with a given emotion label that contain the corresponding MINT tactic.
Model
Overall
DoI
ER
DE
WRM
HL
PEI
SD
Judge: DeepSeek-V4-Pro
Qwen2.5-7B-Instruct
42.53 ± 2.18
9.35 ± 0.49
8.97 ± 0.47
10.09 ± 0.69
9.41 ± 0.19
8.59 ± 0.68
8.85 ± 0.98
7.56 ± 0.91
PERM ( Wang et al., 2026b )
49.07 ± 1.03
10.97 ± 0.15
10.26 ± 0.10
11.98 ± 0.32
10.63 ± 0.32
10.36 ± 0.43
10.90 ± 0.40
9.58 ± 0.43
KARDIA-R1 ( Yuan et al., 2026 )
28.75 ± 1.21
6.00 ± 0.36
6.04 ± 0.23
8.21 ± 0.12
9.61 ± 0.09
6.80 ± 0.46
6.15 ± 0.27
5.52 ± 0.23
RLVER ( Wang et al., 2026c )
44.23 ± 2.23
9.71 ± 0.62
9.39 ± 0.41
10.88 ± 0.35
10.50 ± 0.60
9.36 ± 0.43
9.23 ± 0.47
8.14 ± 0.49
MICA (rep.) ( Zhang et al., 2026a )
51.47 ± 0.58
11.51 ± 0.20
10.92 ± 0.16
11.74 ± 0.19
11.18 ± 0.09
11.22 ± 0.28
9.88 ± 0.49
9.11 ± 0.66
Appendix
Table 6: Complete EQBench3 results under three LLM judges, including Overall and all seven component scores. Values are means ± standard deviations over three runs; all metrics are higher-is-better, and bold and underline mark the best and second-best means.
Model
EPM-Idx
Outcome
Outcome
RDI
Etot
Snet
Judge: DeepSeek-V4-Pro
Qwen2.5-7B-Instruct
15.38 ± 1.06
3.25 ± 0.76
7.5 ± 1.5
0.0 ± 0.0
2.3 ± 0.9
PERM ( Wang et al., 2026b )
17.21 ± 1.17
6.39 ± 2.14
10.6 ± 1.7
0.3 ± 0.4
8.2 ± 4.8
KARDIA-R1 ( Yuan et al., 2026 )
11.42 ± 0.34
0.81 ± 0.34
2.4 ± 1.0
0.0 ± 0.0
0.0 ± 0.0
RLVER ( Wang et al., 2026c )
26.05 ± 3.37
15.34 ± 4.72
17.3 ± 4.5
5.2 ± 2.9
23.5 ± 7.3
Appendix
Table 7: Complete EMPA EPM-Idx and outcome results under three LLM judges.
Model
Efficiency
Stability
Efficiency
Rho
Sproj
Tau
Stability
Rpos
Align
Pen
Judge: DeepSeek-V4-Pro
Qwen2.5-7B-Instruct
25.67 ± 1.15
0.0 ± 0.0
0.0 ± 0.0
77.0 ± 3.4
22.35 ± 2.38
17.7 ± 1.5
22.3 ± 2.4
27.2 ± 3.3
PERM ( Wang et al., 2026b )
23.19 ± 1.07
0.2 ± 0.2
0.2 ± 0.2
69.3 ± 3.3
25.04 ± 2.25
21.3 ± 2.2
25.2 ± 2.3
28.7 ± 2.4
KARDIA-R1 ( Yuan et al., 2026 )
30.91 ± 0.50
0.0 ± 0.0
0.0 ± 0.0
92.7 ± 1.5
12.29 ± 0.76
8.1 ± 0.9
12.9 ± 0.7
15.9 ± 0.7
RLVER ( Wang et al., 2026c )
17.86 ± 3.15
2.0 ± 1.2
2.0 ± 1.1
49.6 ± 8.8
40.87 ± 4.36
37.2 ± 4.5
39.4 ± 4.2
46.0 ± 5.1
Appendix
Table 8: Complete EMPA results under three LLM judges (continued).
Model
Selections
Preference (%) ↑
MICA (rep.) ( Zhang et al., 2026a )
104
23.9
RLVER ( Wang et al., 2026c )
107
24.6
CARE(R)
224
51.5
Appendix
Table 9: Blinded human preference evaluation. fifteen annotators each evaluate 30 randomly assigned contexts. Model identities are hidden and response order is randomly shuffled.