Persona Following Is Not Selective Control: The Neutrality Gap in LLM User Simulation
Organizations: City University of Hong Kong · Stanford University
Abstract
Persona prompting is widely used to construct user simulations with large language models (LLMs), yet it relies on a largely untested assumption: specifying one user attribute should change that attribute alone. We test this assumption and identify a systematic failure of selective control: across all eight black-box LLMs we audit, changing a target attribute also shifts responses on unspecified, non-target attributes. For example, describing a user as more risk-seeking shifts color choices, even though the prompt never mentions color; we term this cross-attribute influence. Semantic, contextual, and internal analyses collectively suggest that models treat a persona prompt as evidence about the user and extend the inferred profile to unspecified preferences, a process we call trait-conditioned completion. We next ask whether explicitly specifying non-target attributes restores selective control. When a non-target attribute is assigned a clear direction, models generally follow the declaration and suppress the target attribute's influence. However, when the same attribute is declared neutral, the target continues to affect choices across all five open-weight checkpoints, even when the model correctly reports the declared state. This disparity, the neutrality gap, demonstrates that successful persona following does not imply selective persona control, which additionally requires keeping non-target attributes stable. We operationalize this distinction with a three-state diagnostic that leaves the non-target attribute unspecified or declares it directional or neutral; because directional tests can be passed by simply following the stated persona, the neutral state reveals failures they miss. In a post hoc analysis of independent items, neutral declarations leave 51-81% of items target-sensitive, against at most 1 of 320 item-pole comparisons under directional ones.
Figures & tables
| Design | Items and split | Models | Readout |
|---|---|---|---|
| Field audit | non-target, target items; -item bank | black-box | samples per item |
| Structured worlds | texts; validation, held-out worlds | open-weight | answer probabilities |
| Independent items | items ( non-target), four pairs | primary | answer probabilities |
| Internal coupling | items, conditions, two item folds | + held-out | seven-point position |
| Comparisons: choice sensitive / stable | ||||||
|---|---|---|---|---|---|---|
| Item set | Check- points | Correct report | Report sens. | Choice sens. | Adequate report | Inadequate report |
| Structured texts | – | – | / of | / | ||
| Independent items | – | – | / of | / | ||
Appendix figures & tables37 assets
Supplementary material from the paper’s appendix.
Appendix
| Claim | Evidence retained here | Main qualification |
|---|---|---|
| C1. Cross-attribute influence | Per-model audit, pole controls, unrelated control attributes, construct coding, and the baseline-clause deletion (App. A.7 ). | Black-box scores are relative to a baseline stating the cautious direction; the coding is author-assigned, but two-model recodings keep the risk–control gap . |
| C2. Trait-conditioned completion | Semantic controls, correlation manipulation, and same-item base–instruct prediction (App. A.8 ). | The arbitrary code is not activation-matched (the scoped invented name is the activation-retaining contrast); base prediction is not a causal estimate of post-training. |
| C3. Partial internal coupling | Erasure design, point estimates, intervals, control checks, and readout and utility limits (App. A.9 ). | Three of four held-out checkpoints replicate; the fourth is uninterpretable. No general-utility or neutral-state mechanism claim. |
| C4. Neutrality gap | Prompt-design selection, held-out and independent-item results, scale decomposition, wording controls, and ground-truth identifiability (App. A.10 ). | Holdout reuses texts (the independent items supply new ones); level splits are post hoc; neutral fidelity is not identified and not needed for the invariance test. |
| C5. Reporting versus preserving | Branched-task design, full per-pair results, frame and coverage limits, and the failed internal positive control (App. A.11 ). | Exploratory task-level dissociation, not a localized internal process. |
| C6. Mitigation boundary | Containment rules and full counts, scope-only controls, the rule-lookup reference, and the dialogue result used in the Conclusion (App. A.12 ). | A supplied direction is not neutrality; the rule reference is not a matched neutral condition. |
| Main text | Appendix term | Defined in |
|---|---|---|
| non-target shift | leakage | App. A.3 , A.9 |
| author-annotated attribute labels | construct scope specification | App. A.3 |
| attention-grabbing, non-default pole | marked pole; markedness | App. A.3 |
| non-target attribute, as a prompt factor | discriminant factor | App. A.3.1 |
| field audit | field track | App. A.3 |
| higher-sample audit (Section 3 ) | dense audit | App. A.2 |
| Design | Materials and comparison | Unit / status |
|---|---|---|
| Headline field audit | non-target style items and risk-axis target items, draws per cell; same-axis pole reversals, draws per cell. | Item; prespecified. |
| Factorial / worked criterion | Full bank: three prompt conditions, four factor cells, non-target and target items, draws; compressed factorial with a stated default pole: six profiles, non-target and target items, draws. | Item / cell; descriptive, not certified (App. A.5 ). |
| Dense audit | items outside the target domain ( aesthetic, communication, ambient, ingestive) and risk items; seven-point responses, draws; ingestive (ambiguous) cells excluded from headlines. | Item; descriptive; coding sensitivity post hoc. |
| Semantic controls | non-target and target items, scored as answer-token log-odds under five wording conditions. | Item; descriptive contrasts. |
| Base and representation bank | seven-point items: non-target, ambiguous, target; condition wordings in the representation program; two fit/evaluation folds. | Item; prespecified program, with logged deviations. |
| Explicit-evidence bank | worlds, per attribute pair, non-target texts; demonstrations; absent, declared, mentioned, or past-choice evidence; no neutral level. | World; exploratory. |
| Symbol | Meaning | Where |
|---|---|---|
| , ; , | target and non-target attribute; their assigned values ( is the declared-neutral level) | Sections 2 , 5 |
| , , | constructed world; non-target or target query; prompt design (also the field-audit design index) | Section 2 , Appendix A.3 |
| , | answer distribution under prompt design ; ground-truth reference | Section 2 |
| , , , | non-target sensitivity; attribute-fidelity error; conservative per-world residual; unassigned answer mass | Section 2 , Appendix A.10 |
| , | demonstrated correlation; signed non-target-option difference between the target levels | Section 3 , Appendix A.10 |
| , , , | semantic log-odds at target level ; shared answer tendency; target-induced margin; logistic sigmoid | Section 5 , Appendix A.10.6 |
| Factorial, no instruction | Compressed factorial | ||||||
|---|---|---|---|---|---|---|---|
| Model | (worst) | ? | ? | ||||
| GPT-5.5 | .046–.093 (1.00) | no | 1.000 | .000 | 1.000 | yes | |
| Kimi-K2.6 | .172–.356 (1.00) | no | 1.000 | .000 | 1.000 | yes | |
| Gemini-3-Flash-Preview | .199–.270 (1.00) | no | .906 | .000 | 1.000 | yes | |
| GLM-5.2 | .068–.166 (1.00) | no | .875 | .000 | 1.000 | yes | |
| Qwen3.5-397B | .087–.155 (1.00) | no | 1.000 | .000 | 1.000 | yes | |
| Persona | Expected pole | Basis | Status | Headline? |
|---|---|---|---|---|
| risk-seeking | marked / vivid / intense | semantic hypothesis, not a construct-scope code | prespecified | yes, with caveat |
| bold / flamboyant | marked / expressive | definitionally aligned | prespecified | yes |
| frugal | plain / sparse / default | anti-excess association | conservative | secondary |
| minimalist | plain / sparse | definitionally aligned | prespecified | yes |
| punctual | no aesthetic direction | unrelated control | prespecified | yes (as null) |
| detail-oriented | no documented markedness mapping | unrelated control | conservative | yes (as null) |
| Trait | Domain | Code | Basis / ambiguity reason | Headline? |
|---|---|---|---|---|
| risk-seeking | investment / lottery | target ( risk) | domain-specific risk attitude ( Weber et al., 2002 ; Blais & Weber, 2006 ) | target |
| risk-seeking | clean aesthetic | non-target | no construct-theoretic path in the scope specification; not a zero-association premise, since Big Five traits predict how central visual design is to consumer choice in a human sample ( Myszkowski & Storme, 2012 ) | yes |
| risk-seeking | communication | non-target | no construct-theoretic path; risk attitude is domain-specific ( Weber et al., 2002 ; Blais & Weber, 2006 ) | yes |
| risk-seeking | ambient audio | non-target (sensitivity) | treated with sensitivity checks; associations between sensation seeking and music reported in humans ( Litle & Zuckerman, 1986 ) | yes |
| risk-seeking | ingestive / travel | ambiguous | stimulation-adjacent covariance plausible ( Zuckerman, 1994 ) | excluded |
| openness | art / design | target ( complex) | aesthetic-openness literature ( McCrae & Costa, 1997 ; Cleridou & Furnham, 2014 ) | positive control |
| Shift on non-target items, by prompt | Pole reversal | |||
| Model | Risk-seeking | Bold (aligned) | Cautious (opposite pole) | plus / minus |
| DeepSeek-V4-Flash | / | |||
| GPT-5.5 | / | |||
| Kimi-K2.6 | / | |||
| Gemini-3-Flash-Preview | / | |||
| GLM-5.2 | / | |||
| Risk-stating baseline | Clause-deleted baseline | Deleted stated | |||
|---|---|---|---|---|---|
| Checkpoint | cautious | cautious | cautious | ||
| Qwen2.5-32B | |||||
| Qwen3-32B | |||||
| OLMo-2-32B | |||||
| Gemma-2-27B | |||||
| Mistral-Small-3.1-24B | |||||
| Condition | On-construct (18 items) | Leakage (94 items) |
|---|---|---|
| Natural label | ||
| Operational paraphrase | ||
| Invented name | ||
| Scoped name | ||
| Arbitrary code |
| Model | |||
|---|---|---|---|
| Mistral | |||
| Qwen3-32B | |||
| OLMo-2-32B | |||
| Gemma-2-27B | |||
| Qwen2.5-32B |
| Model | full | own | markedness | target’s own | target- orthogonal |
|---|---|---|---|---|---|
| Qwen2.5-32B | |||||
| Qwen3-32B | |||||
| Gemma-2-27B ‡ | |||||
| OLMo-2-32B | |||||
| Mistral | |||||
| Llama-3.1-8B-Instruct |
| Model | marked CEP | marked RET | Target-orthogonal CEP | Target-orthogonal RET |
|---|---|---|---|---|
| Qwen2.5-32B | ||||
| Qwen3-32B | ||||
| Gemma-2-27B † | ||||
| OLMo-2-32B | ||||
| Mistral | ||||
| Model | answer CEP | marked answer |
| Audit question | Result | Supported interpretation |
|---|---|---|
| Fixed-setting rank- seven-point check | Passes every control on four; Gemma-2 misses only the side-effect rule ( against ). | |
| Adjacent rank/depth region | Late dependence survives confidence-bound testing. | |
| Held-out seven-point replication | The fourth checkpoint is uninterpretable because its random control fails. | |
| Binary check, transported subspaces | Transport without refitting is not panel-wide; refitting inside each readout defines no pass criterion and is format-dependent and model-heterogeneous (Appendix A.9.3 ), so no count applies to it. | |
| Answer span: non-containment / independence | / | Markedness still removes the leakage with the answer span projected out, but that span alone also removes it. |
| All collateral-utility criteria | Two more checkpoints miss one criterion each; selectivity on audited items is not broad utility preservation. |
| Checkpoint | Seven-point | Binary choice |
|---|---|---|
| Qwen3.5-9B | ||
| Qwen3-32B | ||
| Mistral-Small-3.1-24B | ||
| Gemma-2-27B | ||
| OLMo-2-32B |
| Endpoint contrast | Probability fraction | Margin fraction | Declared median | |
| Model | bare / declared | [95% interval] | [95% interval] | absolute margin (nat) |
| Mistral | 1.3963 / 0.0358 | 0.0256 [0.0109, 0.0443] | 0.0975 [0.0775, 0.1203] | 6.3125 |
| Qwen3 | 1.1406 / 0.0662 | 0.0581 [0.0235, 0.0954] | 0.3058 [0.2600, 0.3540] | 16.0000 |
| OLMo-2 | 1.3077 / 0.0459 | 0.0351 [0.0183, 0.0567] | 0.1965 [0.1688, 0.2264] | 9.3750 |
| Gemma-2 | 0.8951 / 0.0027 | 0.0031 [ 0.0335, 0.0296] | 0.0703 [0.0421, 0.0979] | 10.2500 |
| Qwen2.5 | 1.3537 / 0.0075 | 0.0056 [ 0.0000, 0.0182] | 0.0474 [0.0389, 0.0575] | 33.3750 |
| Bare | Declaration | Selected | Selected declaration | Selected | Selected | |
|---|---|---|---|---|---|---|
| Model | sensitivity | sensitivity | sensitivity | [95% interval] | fidelity error | target effect |
| Mistral | 0.7196 | 0.1033 | 0.0891 | 0.0143 [ 0.0225, 0.0063] | 0.1087 | 0.9830 |
| Qwen3 | 0.5968 | 0.1282 | 0.1205 | 0.0077 [ 0.0124, 0.0035] | 0.1196 | 0.9042 |
| OLMo-2 | 0.6987 | 0.1131 | 0.1106 | 0.0025 [ 0.0071, 0.0021] | 0.1141 | 0.9903 |
| Gemma-2 | 0.7750 | 0.1647 | 0.1653 | 0.0006 [ 0.0097, 0.0112] | 0.1190 | 0.9520 |
| Qwen2.5 | 0.7161 | 0.1436 | 0.0674 | 0.0762 [ 0.0926, 0.0597] | 0.1088 | 0.9971 |
| Declaration | Selected | Selected declaration | Declaration | Selected | Selected | |
|---|---|---|---|---|---|---|
| Model | sensitivity | sensitivity | [95% interval] | fidelity error | fidelity error | target effect |
| Mistral | 0.0043 | 0.0018 | 0.0026 [ 0.0044, 0.0012] | 0.0045 | 0.0030 | 0.9830 |
| Qwen3 | 0.0134 | 0.0189 | 0.0055 [ 0.0002, 0.0131] | 0.0085 | 0.0127 | 0.9042 |
| OLMo-2 | 0.0010 | 0.0042 | 0.0033 [0.0003, 0.0074] | 0.0092 | 0.0076 | 0.9903 |
| Gemma-2 | 0.0018 | 0.0055 | 0.0038 [ 0.0003, 0.0097] | 0.0010 | 0.0029 | 0.9520 |
| Qwen2.5 | 0.0420 | 0.0062 | 0.0359 [ 0.0553, 0.0175] | 0.0210 | 0.0031 | 0.9971 |
| Model | Conservative share | Normalized-TV share [text CI] | texts / | all / no |
|---|---|---|---|---|
| Mistral | / | |||
| Qwen3 | / | |||
| OLMo-2 | / | |||
| Gemma-2 | / | |||
| Qwen2.5 | / |
| Conservative | Normalized-TV | Conservative | Nonzero-z | Extreme | |
|---|---|---|---|---|---|
| Model | union /477 | union /477 | count | union | union |
| Mistral | 431 (90.4%) | 293 (61.4%) | 431 | 45 | 0 |
| Qwen3 | 303 (63.5%) | 279 (58.5%) | 281 | 74 | 46 |
| OLMo-2 | 296 (62.1%) | 296 (62.1%) | 288 | 43 | 8 |
| Gemma-2 | 364 (76.3%) | 364 (76.3%) | 364 | 77 | 8 |
| Qwen2.5 | 184 (38.6%) | 184 (38.6%) | 180 | 13 | 6 |
| Model | Union | |||||
|---|---|---|---|---|---|---|
| Mistral | 0 | 8 | 285 | 35 | 0 | 293 |
| Qwen3 | 0 | 0 | 248 | 18 | 17 | 279 |
| OLMo-2 | 0 | 10 | 288 | 25 | 8 | 296 |
| Gemma-2 | 8 | 13 | 364 | 64 | 0 | 364 |
| Qwen2.5 | 0 | 0 | 180 | 13 | 6 | 184 |
| Sensitivity threshold: neutral / directional share | ||||||||
| Model | ||||||||
| Mistral | / | / | / | / | / | / | / | |
| Qwen3 | / | / | / | / | / | / | / | |
| OLMo-2 | / | / | / | / | / | / | / | |
| Gemma-2 | / | / | / | / | / | / | / | |
| Qwen2.5 | / | / | / | / | / | / | / | |
| Model | exceed. worlds / texts | Tail, normalized TV rate [CI] (lower) | Tail, conservative rate [CI] | all [CI] | no [CI] | Fidelity error, directional [CI] |
|---|---|---|---|---|---|---|
| Mistral | / | ( ) | ||||
| Qwen3 | / | ( ) | ||||
| OLMo-2 | / | ( ) | ||||
| Gemma-2 | / | ( ) | ||||
| Qwen2.5 | / | ( ) |
| Share above | Neutral: items above of | ||||||
|---|---|---|---|---|---|---|---|
| Checkpoint | Bare | Directional | Neutral [95% CI] | Color | Hue | Sound | Layout |
| Qwen3-32B | |||||||
| OLMo-2-32B | |||||||
| Gemma-2-27B | |||||||
| Mistral-Small-3.1-24B | |||||||
| Qwen3.5-9B | |||||||
| Model | Large shifts at | Same sign as bare | Opposite | Toward / away from truth | Argmax flips | Log-odds (nat) | Log-odds (nat) extremes |
|---|---|---|---|---|---|---|---|
| Mistral | 285 | 255 | 0 | 153 / 132 | 165 | 2.96 | 0.34 |
| Qwen3 | 248 | 222 | 0 | 131 / 117 | 230 | 10.66 | 1.31 |
| OLMo-2 | 288 | 277 | 0 | 150 / 138 | 222 | 6.65 | 0.91 |
| Gemma-2 | 364 | 321 | 0 | 181 / 183 | 296 | 10.03 | 0.88 |
| Qwen2.5 | 180 | 162 | 3 | 97 / 83 | 127 | 9.79 | 2.87 |
| Fidelity error by | Mean fidelity error, | Mean sensitivity, | |
|---|---|---|---|
| Model | all levels / directional only | all levels / directional only | |
| Mistral | .003, .008, .492, .032, .003 | .108 / .011 | .081 / .014 |
| Qwen3 | .000, .000, .494, .045, .019 | .112 / .016 | .107 / .013 |
| OLMo-2 | .000, .007, .492, .034, .014 | .109 / .014 | .104 / .013 |
| Gemma-2 | .007, .011, .487, .052, .000 | .111 / .018 | .155 / .035 |
| Qwen2.5 | .000, .000, .499, .012, .006 | .103 / .005 | .061 / .009 |
| Model | Level | mean | mean | median ratio | same sign | mean TV | |
|---|---|---|---|---|---|---|---|
| Mistral | bare | – | – | ||||
| Qwen3 | bare | – | – | ||||
| Model | Pair | bare | dir. / | W0 battery line | W1 no stable tendency | W2 no settled preference | W3 indifferent between poles | W4 W0 first-listed rule |
|---|---|---|---|---|---|---|---|---|
| Qwen3 | risk color | ; | / ; | ; | ; | ; | ; | ; |
| Qwen3 | time sound | ; | / ; – | ; | ; | ; | ; | ; |
| Mistral | risk color | ; | / ; | ; | ; | ; | ; | ; |
| Mistral | time sound | ; | / ; | ; | ; | ; | ; | ; |
| Checkpoint | No instruction | Relevance none | Scope none |
|---|---|---|---|
| risk weight color preference | |||
| Qwen2.5-32B | |||
| Qwen3-32B | |||
| OLMo-2-32B | |||
| Gemma-2-27B | |||
| Mistral-Small-3.1-24B | |||
| Model | exceedances all / / | Normalized-TV tail all / (enter) | Sign kept / of | Sign and magnitude kept / |
|---|---|---|---|---|
| Mistral | / / | / ( ) | / of | / |
| Qwen3 | / / | / ( ) | / of | / |
| OLMo-2 | / / | / ( ) | / of | / |
| Gemma-2 | / / | / ( ) | / of | / |
| Qwen2.5 | / / | / ( ) | / of | / |
| Model | Pair | Ret. | |||||
|---|---|---|---|---|---|---|---|
| (a) Structured texts, six checkpoints | |||||||
| Qwen2.5 | risk color | ||||||
| Qwen2.5 | time sound | ||||||
| Qwen3 | risk color | ||||||
| Qwen3 | time sound | ||||||
| Mistral | risk color | ||||||
| Instruction wording | Contained | Leaks | Over-corrects | Dilutes | Contained share |
|---|---|---|---|---|---|
| Scope ordinary default | |||||
| Task-specific | |||||
| Ordinary default only | |||||
| Scope only | |||||
| Negation-only | |||||
| Domain-label prefix |
| Target question | Non-target question | ||||||
| Rule agreement | |||||||
| Checkpoint | all | Target effect | Rule agreement | Sensitivity | |||
| Mistral | |||||||
| Qwen3 | |||||||
| OLMo-2 | |||||||
| Gemma-2 | |||||||
| Model | Bare | Ratio [CI] | Scoped | Ret. | Repair |
|---|---|---|---|---|---|
| DeepSeek-V4-Flash | fails (t1) | ||||
| GPT-5.5 | fails (t5) | ||||
| Kimi-K2.6 | fails (t1) | ||||
| Gemini-3-Flash-Preview | persists | ||||
| GLM-5.2 | persists | ||||
| Qwen3.5-397B | fails (t1) |
| Model | Provider-facing ID | API provider | Think | Valid records |
|---|---|---|---|---|
| DeepSeek-V4-Flash | deepseek-v4-flash | DeepSeek API | on | |
| GPT-5.5 | gpt-5.5 | commercial gateway A | on | |
| Kimi-K2.6 | kimi-k2.6 | commercial gateway A | off | |
| GLM-5.2 | glm-5.2:cloud | Ollama API | off | |
| Nemotron-3-Super | nemotron-3- super:cloud | Ollama API | off | |
| Qwen3.5-397B | qwen3.5:397b-cloud | Ollama API | off |
| Model | Endpoint type | Std. ratio | CI |
|---|---|---|---|
| GLM-5.2 | open-weight via API | ||
| DeepSeek-V4-Flash | open-weight via API | ||
| GPT-5.5 | proprietary closed API | ||
| MiMo-V2.5-Pro | open-weight via API | ||
| Qwen3.5-397B | open-weight via API | ||
| Gemini-3-Flash-Preview | proprietary closed API |
| Display name | Instruct checkpoint | Base checkpoint |
|---|---|---|
| Qwen3-32B | Qwen/Qwen3-32B | none public at this scale; other families’ bases predict it ( ) |
| Qwen2.5-32B | Qwen/Qwen2.5-32B-Instruct | Qwen/Qwen2.5-32B |
| Mistral-Small-3.1-24B | mistralai/Mistral-Small-3.1-24B-Instruct-2503 | mistralai/Mistral-Small-3.1-24B-Base-2503 |
| Gemma-2-27B | google/gemma-2-27b-it | google/gemma-2-27b |
| OLMo-2-32B | allenai/OLMo-2-0325-32B-Instruct | allenai/OLMo-2-0325-32B |
| Gemma-3-27B | google/gemma-3-27b-it | google/gemma-3-27b-pt |