cs.AISep 28, 2026

Persona Following Is Not Selective Control: The Neutrality Gap in LLM User Simulation

Authors: Jiashen Ren, Wenlin Zhang, Bohan Zhang, Xiaopeng Li, Zichuan Fu, Wanyu Wang, Junyi Li, Xiangyu Zhao

Organizations: City University of Hong Kong · Stanford University

Abstract

Persona prompting is widely used to construct user simulations with large language models (LLMs), yet it relies on a largely untested assumption: specifying one user attribute should change that attribute alone. We test this assumption and identify a systematic failure of selective control: across all eight black-box LLMs we audit, changing a target attribute also shifts responses on unspecified, non-target attributes. For example, describing a user as more risk-seeking shifts color choices, even though the prompt never mentions color; we term this cross-attribute influence. Semantic, contextual, and internal analyses collectively suggest that models treat a persona prompt as evidence about the user and extend the inferred profile to unspecified preferences, a process we call trait-conditioned completion. We next ask whether explicitly specifying non-target attributes restores selective control. When a non-target attribute is assigned a clear direction, models generally follow the declaration and suppress the target attribute's influence. However, when the same attribute is declared neutral, the target continues to affect choices across all five open-weight checkpoints, even when the model correctly reports the declared state. This disparity, the neutrality gap, demonstrates that successful persona following does not imply selective persona control, which additionally requires keeping non-target attributes stable. We operationalize this distinction with a three-state diagnostic that leaves the non-target attribute unspecified or declares it directional or neutral; because directional tests can be passed by simply following the stated persona, the neutral state reveals failures they miss. In a post hoc analysis of independent items, neutral declarations leave 51-81% of items target-sensitive, against at most 1 of 320 item-pole comparisons under directional ones.

Figures & tables

Appendix figures & tables37 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. How Well Do Large Language Models Capture Human Personality?

    May 12, 2026Aanisha Bhattacharyya, Yaman Kumar Singla, Rajiv Ratn Shah +2Artificial Intelligence PersonasPersona Consistency

  2. PrivacySIM: Evaluating LLM Simulation of User Privacy Behavior

    May 12, 2026James Flemings, Murali AnnavaramUser SimulationPersona Consistency

  3. The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study

    May 20, 2026Victoria Lin, Taedong Yun, Maja Matarić +3User SimulationIllusions