cs.AISep 28, 2026

Evolving Support Priorities in Empathetic Reinforcement Learning

Authors: Pengyu Huang, Zhiyuan Han, Wenwen Tong, Hewei Guo, Jiangnan Chen, Sirui Chen, Lewei Lu, Beier Zhu, +1 more

Organizations: University of Science and Technology of China, Hefei, China · SenseTime Research, Shanghai, China · Institute of Artificial Intelligence, Hefei Comprehensive National Science Center

Abstract

We identify a fundamental mismatch in empathetic reinforcement learning: support priorities evolve with the dialogue state, yet existing methods typically optimize predefined reward specifications that remain fixed across turns. To model these evolving support priorities, we organize empathetic support along cognitive, affective, and proactive empathy, and propose Context-Adaptive Rubric Evolution (CARE). At each turn, CARE generates a context-adaptive rubric by adjusting both the weights of these three empathy dimensions and their fine-grained evaluation criteria. The rubric generator is trained with turn-level rubric supervision and human preference data through supervised fine-tuning followed by preference-based reinforcement learning, and then serves as an adaptive reward interface for online empathetic RL. Integrated with both RLVER and MICA, CARE achieves state-of-the-art performance across SentientBench, EQBench3, and EMPA under three independent LLM judges. Notably, on EMPA, CARE improves EPM-Idx over the strongest baseline by at least 13 points under all three judges, including an increase from 28.11 to 83.54 under Gemini-2.5-Pro. Further analyses show that learned rubric priorities systematically vary across dialogue stages and user emotions, demonstrating that CARE adapts what is rewarded as support needs evolve.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue

    Aug 11, 2026Yi Wei, Shuo Jiang, Huaixia Dou +5EmpathySelf-Evolution

  2. Can You Break RLVER? Probing Adversarial Robustness of RL-Trained Empathetic Agents

    May 8, 2026Deeraj S K, Sadhana Devarajan, Krishna Mehra +1EmpathyReinforcement Learning With Verifiable Reward

  3. From Empathy to Personalized Empathy: Adapting Empathetic Strategies to Individual Users

    May 30, 2026Wuqiang Zheng, Chengbing Wang, Yilin Yang +6Empathy