As agentic systems gain commercial popularity, user simulators increasingly serve as measurement instrument for their evaluation. However, the fidelity of simulated users in comparison to real human users is generally low, and typically assessed by costly, subjective LLM judges. In this pilot study, we ask whether fidelity can instead be measured deterministically by treating a user persona sociolinguistically: as a social type that emerges from observable linguistic style, rather than one predicted by labels or descriptions a model must extrapolate into behaviour. We author personas as concrete stylistic rates, which lets us transfer two established, model-free instruments -- authorship-verification stylometry and lexicon-based content analysis -- as fidelity diagnostics. We A/B-test the sociolinguistic schema against a flat descriptive baseline across five task-oriented customer-service agents. Results show that the sociolinguistic schema improves both stylistic adherence and stylometric distinguishability for most of the tested models, with a caveat that persona style fidelity does not necessarily equal persona "naturalness". We argue that a sociolinguistic approach to persona design is a promising path towards more diverse and representative user personas, and that these metrics are most valuable in an error-attribution analysis, localizing where fidelity breaks down. This is a first step towards interventions that move user simulations closer to faithful renderings of diverse and variable linguistic outputs.
Figures & tables
Figure 1: From an identical prompt, the AnxiousInquirer and FrustratedComplainer personas diverge into opposite styles. Du Bois (2007) ’s stance triangle holds that an utterance at once performs three actions: evaluates the object of talk, positions the speaker’s affect and certainty, and (dis)aligns the speaker with the interlocutor. All three are expressed (often subtly) through overt linguistic forms, making stance observable at the surface level. The two turns above invert all three vectors: self-blame, hedges, filler, and an apologetic emoticon (AnxiousInquirer) versus company-blame, a negative evaluative epithet, and a bare imperative (FrustratedComplainer).
archetype
definition
Δ
text
gpt-4o-mini
0.483
0.593
+0.111
gpt-5-mini
0.449
0.636
+0.187
gemini-3.5-flash (low)
0.408
0.423
+0.015
gemini-3.5-flash (high)
0.481
0.482
+0.001
gemini-3.1-pro-preview
0.417
0.414
−0.003
Table 1: Stylometric distinguishability (AUC) across text and voice, mean over five agents.
archetype
definition
Δ
per-persona
gpt-4o-mini
0.194
0.428
+0.235
gpt-5-mini
0.153
0.573
+0.420
gemini-3.5-flash (low)
−0.042
−0.032
+0.010
gemini-3.5-flash (high)
−0.041
0.014
+0.054
gemini-3.1-pro-preview
0.023
0.023
+0.000
Table 2: Stylistic adherence ( τb ) across text and voice, mean over five agents. Voice is averaged over the six non-orthographic features it can realize; all-caps and CMC-acronyms are text-only (§ 3.3.2 ).
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
metric
archetype
no block
with block
Δ
Distinguishability AUC
0.453
0.518
0.652
−0.134
Adherence τb (per-persona)
0.181
0.489
0.619
−0.130
Adherence τb (per-conversation)
0.106
0.379
0.515
−0.136
Appendix
Table 3: Ablation of the behavioural-contracts block, gpt-5-mini .
gpt-4o-mini
gpt-5-mini
voice
Persona
same
diff
gap
same
diff
gap
same
diff
gap
anxious_inquirer
0.618
0.469
+0.149
0.648
0.482
+0.166
0.640
0.531
+0.109
vague_storyteller
0.592
0.460
+0.131
0.614
0.478
+0.135
0.677
0.529
+0.148
methodical_researcher
0.512
0.439
+0.073
0.582
0.462
+0.120
0.535
0.484
+0.051
technophobe
0.465
0.439
+0.026
0.528
0.442
+0.087
0.575
0.501
+0.074
social_chatterbox
0.545
0.449
+0.096
0.538
0.458
+0.080
0.610
0.515
+0.096
Appendix
Table 4: Per-persona separability ( definition arm), OpenAI models, mean over five agents.
gemini-3.5-flash (low)
gemini-3.1-pro-preview
gemini-3.5-flash (high)
gemma-4-31b-it
Persona
same
diff
gap
same
diff
gap
same
diff
gap
same
diff
gap
anxious_inquirer
0.323
0.357
−0.034
0.327
0.372
−0.045
0.238
0.250
−0.012
0.189
0.200
−0.012
vague_storyteller
0.303
0.356
−0.053
0.347
0.382
−0.035
0.221
0.244
−0.022
0.173
0.194
−0.020
methodical_researcher
0.318
0.360
−0.042
0.327
0.373
−0.045
0.229
0.247
−0.018
0.178
0.197
−0.019
technophobe
0.347
0.376
−0.028
0.332
0.377
−0.045
0.243
0.254
−0.011
0.195
0.203
−0.008
social_chatterbox
0.307
0.361
−0.054
0.308
0.370
−0.062
0.246
0.254
−0.008
0.169
0.193
−0.023
Appendix
Table 5: Per-persona separability ( definition arm), Gemini and Gemma models, mean over five agents.
claude-sonnet-5
claude-haiku-4-5
claude-opus-5
Persona
same
diff
gap
same
diff
gap
same
diff
gap
anxious_inquirer
0.540
0.387
+0.153
0.637
0.490
+0.147
0.595
0.381
+0.214
vague_storyteller
0.438
0.379
+0.059
0.663
0.502
+0.161
0.469
0.363
+0.106
methodical_researcher
0.401
0.363
+0.037
0.516
0.448
+0.068
0.476
0.364
+0.112
technophobe
0.506
0.378
+0.128
0.597
0.489
+0.107
0.543
0.379
+0.165
social_chatterbox
0.416
0.362
+0.054
0.661
0.496
+0.165
0.475
0.342
+0.134
Appendix
Table 6: Per-persona separability ( definition arm), Claude models, mean over five agents.
gpt-4o-mini
gpt-5-mini
voice
Feature
arch.
def.
Δ
arch.
def.
Δ
arch.
def.
Δ
CMC-acronym rate
0.000
0.305
+0.305
0.000
0.838
+0.838
N/A
N/A
N/A
filler rate
0.121
0.654
+0.533
0.033
0.653
+0.620
0.148
0.567
+0.418
all-caps rate
−0.238
0.154
+0.393
−0.126
0.493
+0.619
N/A
N/A
N/A
hedging rate
0.203
0.518
+0.315
0.071
0.680
+0.608
0.119
0.587
+0.468
message length
0.652
0.668
+0.016
0.338
0.734
+0.396
0.619
0.751
+0.132
Appendix
Table 7: Per-feature adherence ( τb , per-persona), OpenAI models, mean over five agents.
gemini-3.5-flash (low)
gemini-3.1-pro-preview
gemini-3.5-flash (high)
gemma-4-31b-it
Feature
arch.
def.
Δ
arch.
def.
Δ
arch.
def.
Δ
arch.
def.
Δ
CMC-acronym rate
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
filler rate
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
all-caps rate
0.054
−0.199
−0.253
0.199
−0.011
−0.210
0.577
0.157
−0.420
−0.157
−0.262
−0.105
hedging rate
0.080
−0.035
−0.115
0.154
−0.120
−0.274
−0.276
0.276
+0.552
−0.197
0.267
+0.464
message length
−0.173
0.008
+0.182
−0.142
0.299
+0.441
−0.619
−0.041
+0.578
−0.371
−0.041
+0.330
Appendix
Table 8: Per-feature adherence ( τb , per-persona), Gemini and Gemma models, mean over five agents.
claude-sonnet-5
claude-haiku-4-5
claude-opus-5
Feature
arch.
def.
Δ
arch.
def.
Δ
arch.
def.
Δ
CMC-acronym rate
0.189
0.564
+0.375
0.284
0.631
+0.347
−0.032
0.556
+0.588
filler rate
0.593
0.577
−0.015
0.580
0.723
+0.144
0.087
0.548
+0.461
all-caps rate
0.072
0.367
+0.295
−0.156
0.573
+0.729
0.106
0.533
+0.427
hedging rate
0.235
0.652
+0.418
0.515
0.782
+0.267
0.080
0.767
+0.687
message length
0.668
0.635
−0.033
0.371
0.619
+0.247
0.437
0.833
+0.396
Appendix
Table 9: Per-feature adherence ( τb , per-persona), Claude models, mean over five agents.
Large Language Model (LLM) agents are increasingly deployed in settings where they interact with diverse users, including those who are unclear, impatient, or reluctant to share information. However, collecting real interaction data at scale remains expensive. The field has turned to LLM-based \emph{user simulators} as stand-ins, but these simulators inherit the behavior of their underlying models: cooperative and homogeneous. As a result, agents that appear strong in simulation often fail in real human interactions. To narrow this gap, we introduce Persona Policies (PPol), a plug-and-play control layer that induces realistic behavioral variation in user simulators while preserving original task goals. Rather than hand-crafting personas, we employ an evolutionary coding agent to discover persona generation programs optimized for human-likeness and behavioral coverage over real user conversations. The evolved program generates diverse, human-like personas for any task in the domain. Across 4 benchmarks--including τ2-bench Retail and Airline, ColBench, and WildChat--evolved PPol yield 28-72% absolute gains in fitness score over the baseline simulator. In blinded evaluations, annotators judged PPol users as 'human' 80.4% of the time, nearly 2x more than the baseline simulators. Training agents with PPol also improves real-world performance: our user study with live human-agent interactions showed that fine-tuning with our method boosted task success by +23% over default baselines. PPol thus offers a novel approach to strengthen simulator-based evaluation and training without changing underlying tasks.
Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks score only the agent and do not directly measure whether the user correctly executed its assigned role. We introduce UserProxyBench, an evaluation layer over the tau-bench family, and the User Fidelity Score (UFS), which measures adherence to the benchmark's private user instructions using task-grounded rubric criteria scored independently of agent success. Holding the agent fixed at GPT-5.5 and varying only the user proxy across 375 enterprise tasks changes mean task reward by 15.2 points, while 24.4% of successful episodes contain a user-specification violation. The dominant failure is premature disclosure: users provide information before it is requested. This behavior has little effect on task reward, yet among successful episodes it causes the agent to make 1.06 fewer tool calls on average, changing the interaction being evaluated while preserving the reward. Finally, across seven proxies we identify an empirical cost-fidelity frontier, enabling practitioners to select the least expensive simulator that satisfies a required fidelity level.
Evaluating tool-augmented LLM agents requires diverse, realistic user inputs yet most evaluation frameworks use flat role descriptions ("you are an angry customer") that produce near-identical conversations regardless of the underlying scenario. In this paper, we propose a three-tier persona vector with 23 operationalized dimensions: 6 categorical demographics (jurisdiction, age, channel, device, language proficiency, time availability), 12 continuous behavioral traits (patience, assertiveness, digital literacy, etc.) sampled with Gaussian noise around curated profile base vectors, and 5 continuous emotional states (frustration, anxiety, trust, confidence, stress) that shift in response to scenario context. Orthogonal to the persona, a 4-level query-complexity overlay controls utterance phrasing from direct to deliberately vague. We evaluate the persona model inside a synthetic data generation pipeline across 64,698 multi-turn conversations spanning 8 named profiles and 3 production corpora. Key findings: (i) a 15.8 percentage-point spread in agent goal-achievement across personas confirms trait vectors produce measurably different user behavior; (ii) the same persona behaves differently across scenarios due to scenario-reactive emotional state shifts, validating the scenario-reactive design; (iii) domain-specific projects show persona sensitivity on booking-flow compliance (~15-20 percentage points gap between tier-aware and pressure-test personas), demonstrating the model faithfully reproduces real-world difficulty distributions; (iv) seven rule-described trait correlations produce auditable co-occurrence patterns without requiring learned covariance matrices. The persona model is fully specified for reproduction.