As voice agents gain more popularity commercially, the user simulators used to evaluate the deployed agents are also being developed to include more realistic, variable, and diverse speech naturalness behaviors -- disfluency, interruption and backchanneling. The quality of the user simulator directly affects the validity of agent evaluation results. However, we find that most studies so far have not examined in detail whether the intended configuration for these behaviors is realized in the simulation. In this study, we audit the realized naturalness behaviors of tau-Voice, our own LLM-based prompting approach across three models, and our rule-based injection algorithm for disfluency, interruption, and backchanneling. We find that prompting for these behaviors is unreliable and produces speech inconsistent with the instructions, placed and distributed less naturally than the instruction implies. In contrast, our rule-based, model-free algorithm produces controllable and diverse naturalness behaviors more aligned with natural speech. Our results suggest that LLMs not purpose-trained for user simulation are not sufficient on their own to represent authentic user behavior, and that linguistically informed deterministic approaches or specialized models are needed to close the gap; auditing and reporting realized naturalness behaviors, rather than configured settings, is what makes that gap visible.
Figures & tables
Fillers / 1k
Repetitions / 1k
Turn-Initial Fillers (%)
Condition
4o-mini
gpt-5
Gemini
4o-mini
gpt-5
Gemini
4o-mini
gpt-5
Gemini
Uninstructed Baseline
0.00
0.55
0.00
0.00
0.28
0.49
—
0.0
—
“occasionally”
19.29
24.97
37.11
0.79
1.31
8.37
67.3
17.9
52.6
“roughly half”
33.22
75.42
64.05
1.51
6.21
22.87
64.8
28.2
56.8
“most sentences”
34.72
73.67
75.28
2.96
6.16
32.66
62.8
37.6
60.3
τ -Voice Prompt
12.89
12.33
11.74
0.81
0.00
0.82
71.9
35.7
60.5
Table 1: Disfluency generation metrics across prompt directives and injection ( n=299 paired dialogue positions per model and condition). Filler and repetition rates are reported per 1,000 words. Turn-initial share represents the percentage of filler particles occurring on the first word of the turn. Pre-TTS injection operates directly on uninstructed model outputs, producing identical text modifications across models.
System Configuration
Opportunities
Actions Taken
% of Opp.
Agent Turns
Per Turn
Engine Interruption ( β )
β=0.0
230
0
0.0%
118
0.000
β=0.3
244
36
14.8%
153
0.235
β=0.7
246
44
17.9%
153
0.288
β=1.0
194
57
29.4%
106
0.538
Engine Backchannel ( γ )
Table 2: Realized interruption and backchannel rates against configured settings across 20 paired scenarios per condition, evaluated against qualifying mid-turn pause opportunities and agent turns.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Condition
Fill.
Rep.
Comb.
95% CI
Init.
Any
recorded live run
0.41
0.00
0.41
[0.0, 1.3]
1.000
0.3%
gpt-4o-mini
none
0.00
0.00
0.00
[0.0, 0.0]
—
0.0%
occasionally
19.29
0.79
20.08
[14.3, 26.1]
0.673 [0.51, 0.83]
16.1%
roughly half
33.22
1.51
34.73
[28.9, 40.6]
0.648 [0.54, 0.75]
28.8%
most sentences
34.72
2.96
37.68
[30.2, 45.1]
0.628 [0.55, 0.72]
29.1%
τ -Voice
12.89
0.81
13.69
[9.2, 18.5]
0.719 [0.56, 0.89]
11.4%
Appendix
Table 3: All conditions, all models, n=299 paired positions. Rates per 1000 words. “Comb.” is filler particles plus repetitions. Intervals on “Comb.” and “Init.” are 95% bootstrap over conversations, 2000 resamples, since positions from one conversation share a persona and a history. “Init.” is the fraction of filler particles at the first word of the turn among all fillers; “Any” is the fraction of turns carrying at least one disfluency. Counting treats an em dash as a word separator, since an unspaced dash otherwise fuses a filler to its neighbor, and counts three repetition forms: adjacent (“the- the”), filler-separated (“I, uh, I”), and two-word phrase (“are you— are you”). Injection is model-independent: it edits recorded replies and makes no call.
Condition
gpt-4o-mini
gpt-5
Gemini 3 Pro
occasionally
0.95 [0.84, 1.07]
0.79 [0.69, 0.89]
0.61 [0.53, 0.70]
roughly half
0.83 [0.74, 0.92]
0.64 [0.55, 0.73]
0.63 [0.53, 0.73]
most sentences
1.00 [0.85, 1.15]
0.60 [0.50, 0.72]
0.66 [0.58, 0.74]
τ -Voice
0.89 [0.85, 0.93]
0.96 [0.86, 1.09]
0.89 [0.81, 1.00]
ablated
0.91 [0.83, 1.03]
0.89 [0.82, 1.01]
0.86 [0.79, 0.95]
Injection (model-independent)
Appendix
Table 4: Variance-to-mean ratio of the per-turn filler-plus-repetition count, n=299 turns per row, with a 95% bootstrap interval over conversations, 2000 resamples. An interval excluding 1.0 separates the condition from independent placement at a constant rate.
Condition
Metric
Pairs
Nonzero
Mean Δ
95% CI
W
p
β=0.3
interruptions
20
18
+1.80
[1.25, 2.45]
0
1.5×10−4
β=0.7
interruptions
20
18
+2.20
[1.40, 3.05]
0
1.6×10−4
β=1.0
interruptions
20
18
+2.85
[1.85, 3.85]
0
1.8×10−4
γ=0.3
backchannels
20
15
+3.90
[2.30, 5.85]
0
6.4×10−4
γ=0.7
backchannels
20
14
+4.70
[2.85, 6.75]
0
9.4×10−4
γ=1.0
backchannels
20
15
+5.15
[2.55, 8.15]
0
6.0×10−4
Appendix
Table 5: Paired comparison against baseline, 20 scenarios per condition. p values are unadjusted; under Holm correction across the six treatment tests the largest becomes 1.8×10−3 .
Comparison
τb
p
Injection setting α
0.275
3.8×10−31
gpt-4o-mini , clause strength
0.254
1.1×10−23
gpt-5 , clause strength
0.520
2.1×10−108
Gemini 3 Pro, clause strength
0.599
1.1×10−150
Appendix
Table 6: Trend tests, Kendall’s τb between configured strength and realized filler-plus-repetition rate.
LLM-based user simulators are increasingly used to evaluate autonomous agents at scale, in place of costly human evaluations. Despite this promise, these simulators exhibit "assistant bias," a tendency to cooperate and pursue task goals. They rarely reproduce the frustration or disengagement that real users exhibit, compromising evaluation validity. Prior work outlines that this bias is baked in during model training, which role-playing prompts fail to override. We analyze this bias from model activations, extracting a user role vector by contrasting how the model represents user versus assistant perspectives on the same dialogue. We observe two findings: (i) the user direction is identifiable in activations, elicits user-like behaviors, and captures characteristics distinct from assistant traits; and (ii) although user-role activation associates with simulation realism and steering strengthens it, it can exaggerate user behaviors and override individual user profiles. Together, our findings provide a representation-level analysis of LLM user simulators, confirming that assistant bias is structurally identifiable and that user behavior can be directionally analyzed.
Daeheon Jeong, Yoonjoo Lee, Eugene Choi +2
1KAIST · University of Michigan · 3Seoul National University +2
Using offline datasets to evaluate conversational agents often fails to cover rare scenarios or to support testing new policies. This has motivated the use of controllable user simulators for targeted, counterfactual evaluation, typically implemented by prompting or fine-tuning large language models. In this work, we formalize controllable simulation as a causal inference problem. By bridging natural language evaluation with off-policy evaluation methodology, we show that the standard practice of training simulators via supervised fine-tuning on post-hoc trajectory labels yields a structurally biased model. Specifically, these labels are inextricably coupled to the data-generating behavior policy, injecting a look-ahead bias that breaks causal consistency. Furthermore, we prove that under policy shift this failure causes the variance of evaluation metrics to explode geometrically, a phenomenon we term controllability collapse. To restore causal consistency, we establish theoretical conditions for accurate simulation and propose practical training mitigations: a priori controls, step-wise dynamic controls, and direct policy-conditioned learning. Empirical evaluation confirms that while standard global controls distort conversational distributions and collapse behavioral diversity, our causally grounded simulators eliminate look-ahead bias, preserve natural variance, and exhibit robust zero-shot generalization to unseen agent behaviors.
User simulators are widely used as scalable environments for training and evaluating interactive assistants. Generating the next user turn is inherently one-to-many: the same profile and dialogue context may support multiple plausible continuations with different local interaction intents. A fluent response may therefore advance the dialogue through an inappropriate intent, such as acceptance rather than repair. Our key insight is that controllable user simulation should separate which local interaction intent the next user turn should realize from how that intent is expressed in language. We introduce UserIDA (User Intent-Directive Alignment), which exposes interaction intent as an explicit per-turn directive. UserIDA defines a six-way intent interface, learns directive-conditioned generation through supervised fine-tuning, and uses intent-calibrated policy optimization during group-based reinforcement learning. The reward preserves composite response quality while ensuring that intent-violating candidates rank below compliant alternatives in mixed groups. On LMSYS-USP, UserIDA achieves 86.6% intent accuracy, outperforming the strongest dedicated user-simulator baseline by 24.3 percentage points while improving semantic and stylistic similarity. In within-context interventions, it realizes at least four of the six target intents in 91.7% of evaluated dialogue states, compared with 22.9% for the strongest external baseline. These results establish per-turn intent control as a complementary dimension to response fidelity in user simulation.
Bo Wang, Ruixing Zhang, Yunqi Liu +4
the State Key Laboratory of Complex and Critical Software Environment, Beihang University