eess.ASSep 1, 2026

Probing Warmth-Mediated Harm in Speech-Enabled LLMs for Mental-Health Conversations

Authors: Eugenia Kim, Bolor-Erdene Jagdagdorj, Dina Pekelis, Leah Zulas, Amanda Minnich

Organizations: Microsoft, Redmond, WA, USA

Abstract

Audio LLM benchmarks measure understanding and dialogue quality, not whether speech-enabled models respond with relational warmth when a vulnerable user discloses a mental-health concern. We introduce a 7-turn scripted-disclosure probe grounded in WHO mental-health clinical guidelines, with each script run on the same model (Azure OpenAI gpt-realtime) in both audio and text-only conditions, and acoustic-prosody analysis of the generated speech. Across 532 responses we identify two audio-specific patterns transcript-only evaluation would miss: at the elicitation turn the model's voice gets shorter, faster, lower-pitched, and quieter rather than warmer (p < .001 for five of seven acoustic features), and the modality gap on relational acceptance, small in aggregate, concentrates in the highest-stakes self-harm/suicide scripts. A two-rater listener study corroborates that perceived warmth is concentrated at specific turns and on bereavement disclosures. Together these patterns indicate that auditing speech-enabled models in mental-health contexts requires evaluating the combined audio-and-text experience the user encounters, not the transcript in isolation. We release the protocol, scoring pipeline, and scripts as a starting point for evaluating speech-enabled models in mental-health contexts.

Explore similar work

CardsList
  1. Can We Trust LLMs for Mental Health Screening? Consistency, ASR Robustness, and Evidence Faithfulness

    May 10, 2026Erfan Loweimi, Sofia de la Fuente Garcia, Samira Loveymi +2HealthcareMental Health

  2. LLUMI: Improving LLM Writing Assistance for Mental Health Support with Online Community Feedback

    May 28, 2026Jiwon Kim, Maya Ajit, Sherry Gong +4Emotional Support ConversationPrivacy-Preserving Language Models

  3. HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

    Aug 25, 2026Matthew Flathers, Phuong Anh Nguyen, Jill Noorily +7Clinical Language Model EvaluationLLM Evaluation