cs.CLJul 1, 2026

Persona Non Grata: LLM Persona-Driven Generations in MCQA are Unstable in Distinct Dimensions

Authors: César Guerra-SolanoXiang Lorraine Li

Organizations: Department of Computer Science, University of Pittsburgh

Abstract

Persona-driven generations (PDGs) have seen prolific use in research and industry applications, where a large language model (LLM) takes on a 'persona' while completing some task. While persona expressed through free-form text (like dialogue) has substantial work investigating stability or consistency, relatively, persona expressed in non-text-heavy outputs (like in multiple-choice question answering, or MCQA) is often overlooked. We work to address this gap, seeking to understand the instability of LLM PDGs in MCQA tasks. We develop three metrics investigating the performance, outcome, and question correctness stability, evaluating three distinct dimensions. Using these metrics, we find that instability varies consistently between model families and model size, and across question domains, with math/commonsense questions leading to greater instability. We also find task prompt format introduces more prediction instability than other hyperparameters, like temperature. Finally, we find that instability is related to task accuracy, and using our instability metrics, find different experimental settings that result in different best and worst personas for tasks, despite their similarity. This reveals the importance of checking hyperparameter instability in PDGs.

Explore similar work

Apr 8, 2026cs.CL

Persona Matters: Effects of Activation Steering on Short Answer Generation and Scoring

Activation-based steering enables inference-time personalization of large language models, but its effects in educational applications are not well understood. We study activation-based persona vectors representing seven character traits in short-answer generation and automated scoring on the ASAP-SAS benchmark, across three language models spanning dense and mixture-of-experts architectures. Persona steering lowers answer quality overall, with much larger effects on open-ended English Language Arts (ELA) prompts than on factual science prompts. Interpretive and argumentative tasks are particularly sensitive, showing up to 11×\times larger degradation. On the scoring side, we observe predictable valence-aligned calibration shifts: evil'' and impolite'' scorers grade more harshly, while good'' and optimistic'' scorers grade more leniently. ELA tasks are 2.5-3×\times more susceptible to scorer personalization than science tasks, and the mixture-of-experts model shows roughly 6×\times larger calibration shifts than the dense models. To our knowledge, this is the first study to systematically examine the effects of activation-steered persona traits in educational generation and scoring. Our findings highlight the need for task- and architecture-aware calibration when deploying personalized models in educational settings.
Yongchao Wu, Aron Henriksson
May 16, 2026cs.CL

Evaluation Drift in LLM Personality Induction: Are We Moving the Goalpost?

Can large language models reliably express a human-like personality, or are they merely mimicking surface cues without a stable underlying profile? To investigate this, we induce personality in LLMs by fine-tuning them on the long-form essays, where each essay is associated with a target Big Five personality profile. We then evaluate the stability and fidelity of the induced personality using the IPIP-NEO questionnaire. Specifically, we ask: (i) does post-training (SFT, DPO, ORPO) stabilize questionnaire scores under prompt rephrasings, and (ii) can it induce target Big Five profiles from unguided essays? Our results demonstrate that fine-tuning consistently reduces variance in questionnaire responses across five models, directly mitigating the evaluation fragility reported in pre-trained models. However, this newfound stability reveals a more fundamental limitation: accuracy on the full five-dimensional profile remains near chance, even when single-trait scores improve. This indicates that unguided essays lack the cues needed for faithful personality expression. We therefore argue for scenario-grounded datasets or interactive elicitation that accumulates test-aligned evidence over time.
Prateek Rajput, Yewei Song, Iyiola E. Olatunji +2
Sep 12, 2025cs.CL

Human Psychometric Questionnaires Mischaracterize LLM Behavior

We examine whether human psychometric questionnaires can serve as reliable tools for characterizing and predicting LLM behavior in everyday user interactions. We analyze eight open-source LLMs by comparing their value and personality profiles derived from two different methods: Likert self-reports on established questionnaires (PVQ-40/21 and BFI-44/10) and generation probabilities over value-laden responses to everyday user queries. The two profiles diverge substantially. Within-construct item consistency, often cited as evidence of stable LLM dispositions, disappears in generation probabilities. We find that established questionnaire items contain explicit lexical cues that allow models to recognize the target construct and respond in alignment-consistent, socially desirable ways, whereas realistic user queries contain far less recognizable cues. In addition, demographic persona prompts shift models' responses to human questionnaires in ways consistent with real human patterns, but no such shifts appear in the generation setting, highlighting that human questionnaires overestimate LLMs' ability to faithfully reproduce expected psychological traits when role-playing demographic personas. Overall, our study indicates that questionnaire scores alone should not be treated as evidence of LLMs' response tendencies in realistic user interactions, and supports generation-probability profiling with ecologically valid items as a complementary behavioral measure.
Woojung Song, Dongmin Choi, Yoonah Park +3