cs.AIAug 29, 2026

Benevolent Bias in Multi-Turn Human-Agent Dialogue

Authors: Qianqi Liu, Jin Huang, Fethiye Irmak Dogan, Hatice Gunes

Organizations: University of Cambridge

Abstract

Bias in human-agent interaction can manifest not only through hostile language but also as benevolent bias, whereby unequal treatment hides behind a warm, positive tone. To make it detectable, we operationalise benevolent bias along two dimensions, tone and treatment, yielding three classes: neutral support, overt bias, and benevolent bias. Building on these definitions, we construct BENEVDIAL, a class-balanced corpus of 362,880 multi-turn support dialogues spanning user and agent demographics, roles, and generators, to support controlled evaluation. We then test two detector families on it: off-the-shelf safety detectors and prompted large language model (LLM) judges. Our findings reveal a notable detection gap: off-the-shelf detectors reliably flag overt bias yet largely fail to identify benevolent bias. LLM judges improve sensitivity when guided by explicit detection criteria, but this comes at the cost of increased misclassification of neutral supportive statements as benevolent bias, a tendency that is further exacerbated by the presence of demographic context. These findings suggest that fair monitoring of human-agent dialogue must look beyond surface cues to whether the agent's treatment is disparate.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants

    Sep 10, 2025Benjamin Sturgeon, Daniel Samuelson, Jacob Haimes +1Human-Ai InteractionHuman Agency

  2. Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI Dialogue

    Date pendingWilliam Hager, Ishika Rathi, Masum Hasan +1Human-Ai InteractionConversational Artificial Intelligence

  3. AI Safety via Debate is Compromised by Cognitive Biases

    Oct 4, 2026Gefei Liu, Sonya Rashkovan, Sophia Lloyd George +2Large Language Model Decisions