cs.CYMay 12, 2026

Simulating Eating Disorder Patients with LLMs: Evaluating Psychological Persona Stability in Multi-Turn Conversations

Authors: Jennifer HaaseJana Gonnermann-MüllerSee Heng YimNicolas LeinsJan MendlingSebastian Pokutta

Organizations: Weizenbaum Institute, Berlin, Germany · HU Berlin, Berlin, Germany · Zuse Institute Berlin, Berlin, Germany · Department of Psychology, University of Hong Kong, Hong Kong

Abstract

Large language model (LLM)-based simulations of clinical patients are increasingly used for research and training, yet their validity requires persona stability: coherent maintenance of an assigned psychological profile across and within conversations. We evaluate this prerequisite using eating disorder personas grounded in five published case vignettes, a dual-assessment framework (self-report + independent observer ratings), and validated psychometric instruments (EDE-Q) with known ground-truth scores. Across six LLMs and two experiments (between-conversation stability (Exp. I) and within-conversation stability (Exp. II)), we find that LLMs are paradoxically too stable and too inaccurate: variability is negligible, yet all models systematically overshoot ground-truth severity by 12-30% of the scale range (0.7-1.8 points on a 0-6 scale). The mechanism is selective stereotyping: models differentiate cases on behavioural items (dietary restraint) but maximise cognitive-affective items (body dissatisfaction, weight preoccupation) at ceiling regardless of case severity. Additional conversational context does not improve accuracy; it compounds the overshoot. LLMs can portray severe eating pathology but lack a representation of moderate clinical presentations, a "missing middle".

Explore similar work

CardsList
  1. How Well Do Large Language Models Capture Human Personality?

    May 12, 2026Aanisha Bhattacharyya, Yaman Kumar Singla, Rajiv Ratn Shah +2PersonaPersonality