As voice agents gain more popularity commercially, the user simulators used to evaluate the deployed agents are also being developed to include more realistic, variable, and diverse speech naturalness behaviors -- disfluency, interruption and backchanneling. The quality of the user simulator directly affects the validity of agent evaluation results. However, we find that most studies so far have not examined in detail whether the intended configuration for these behaviors is realized in the simulation. In this study, we audit the realized naturalness behaviors of tau-Voice, our own LLM-based prompting approach across three models, and our rule-based injection algorithm for disfluency, interruption, and backchanneling. We find that prompting for these behaviors is unreliable and produces speech inconsistent with the instructions, placed and distributed less naturally than the instruction implies. In contrast, our rule-based, model-free algorithm produces controllable and diverse naturalness behaviors more aligned with natural speech. Our results suggest that LLMs not purpose-trained for user simulation are not sufficient on their own to represent authentic user behavior, and that linguistically informed deterministic approaches or specialized models are needed to close the gap; auditing and reporting realized naturalness behaviors, rather than configured settings, is what makes that gap visible.
Figures & tables
Fillers / 1k
Repetitions / 1k
Turn-Initial Fillers (%)
Condition
4o-mini
gpt-5
Gemini
4o-mini
gpt-5
Gemini
4o-mini
gpt-5
Gemini
Uninstructed Baseline
0.00
0.55
0.00
0.00
0.28
0.49
—
0.0
—
“occasionally”
19.29
24.97
37.11
0.79
1.31
8.37
67.3
17.9
52.6
“roughly half”
33.22
75.42
64.05
1.51
6.21
22.87
64.8
28.2
56.8
“most sentences”
34.72
73.67
75.28
2.96
6.16
32.66
62.8
37.6
60.3
τ -Voice Prompt
12.89
12.33
11.74
0.81
0.00
0.82
71.9
35.7
60.5
Table 1: Disfluency generation metrics across prompt directives and injection ( n=299 paired dialogue positions per model and condition). Filler and repetition rates are reported per 1,000 words. Turn-initial share represents the percentage of filler particles occurring on the first word of the turn. Pre-TTS injection operates directly on uninstructed model outputs, producing identical text modifications across models.
System Configuration
Opportunities
Actions Taken
% of Opp.
Agent Turns
Per Turn
Engine Interruption ( β )
β=0.0
230
0
0.0%
118
0.000
β=0.3
244
36
14.8%
153
0.235
β=0.7
246
44
17.9%
153
0.288
β=1.0
194
57
29.4%
106
0.538
Engine Backchannel ( γ )
Table 2: Realized interruption and backchannel rates against configured settings across 20 paired scenarios per condition, evaluated against qualifying mid-turn pause opportunities and agent turns.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Condition
Fill.
Rep.
Comb.
95% CI
Init.
Any
recorded live run
0.41
0.00
0.41
[0.0, 1.3]
1.000
0.3%
gpt-4o-mini
none
0.00
0.00
0.00
[0.0, 0.0]
—
0.0%
occasionally
19.29
0.79
20.08
[14.3, 26.1]
0.673 [0.51, 0.83]
16.1%
roughly half
33.22
1.51
34.73
[28.9, 40.6]
0.648 [0.54, 0.75]
28.8%
most sentences
34.72
2.96
37.68
[30.2, 45.1]
0.628 [0.55, 0.72]
29.1%
τ -Voice
12.89
0.81
13.69
[9.2, 18.5]
0.719 [0.56, 0.89]
11.4%
Appendix
Table 3: All conditions, all models, n=299 paired positions. Rates per 1000 words. “Comb.” is filler particles plus repetitions. Intervals on “Comb.” and “Init.” are 95% bootstrap over conversations, 2000 resamples, since positions from one conversation share a persona and a history. “Init.” is the fraction of filler particles at the first word of the turn among all fillers; “Any” is the fraction of turns carrying at least one disfluency. Counting treats an em dash as a word separator, since an unspaced dash otherwise fuses a filler to its neighbor, and counts three repetition forms: adjacent (“the- the”), filler-separated (“I, uh, I”), and two-word phrase (“are you— are you”). Injection is model-independent: it edits recorded replies and makes no call.
Condition
gpt-4o-mini
gpt-5
Gemini 3 Pro
occasionally
0.95 [0.84, 1.07]
0.79 [0.69, 0.89]
0.61 [0.53, 0.70]
roughly half
0.83 [0.74, 0.92]
0.64 [0.55, 0.73]
0.63 [0.53, 0.73]
most sentences
1.00 [0.85, 1.15]
0.60 [0.50, 0.72]
0.66 [0.58, 0.74]
τ -Voice
0.89 [0.85, 0.93]
0.96 [0.86, 1.09]
0.89 [0.81, 1.00]
ablated
0.91 [0.83, 1.03]
0.89 [0.82, 1.01]
0.86 [0.79, 0.95]
Injection (model-independent)
Appendix
Table 4: Variance-to-mean ratio of the per-turn filler-plus-repetition count, n=299 turns per row, with a 95% bootstrap interval over conversations, 2000 resamples. An interval excluding 1.0 separates the condition from independent placement at a constant rate.
Condition
Metric
Pairs
Nonzero
Mean Δ
95% CI
W
p
β=0.3
interruptions
20
18
+1.80
[1.25, 2.45]
0
1.5×10−4
β=0.7
interruptions
20
18
+2.20
[1.40, 3.05]
0
1.6×10−4
β=1.0
interruptions
20
18
+2.85
[1.85, 3.85]
0
1.8×10−4
γ=0.3
backchannels
20
15
+3.90
[2.30, 5.85]
0
6.4×10−4
γ=0.7
backchannels
20
14
+4.70
[2.85, 6.75]
0
9.4×10−4
γ=1.0
backchannels
20
15
+5.15
[2.55, 8.15]
0
6.0×10−4
Appendix
Table 5: Paired comparison against baseline, 20 scenarios per condition. p values are unadjusted; under Holm correction across the six treatment tests the largest becomes 1.8×10−3 .
Comparison
τb
p
Injection setting α
0.275
3.8×10−31
gpt-4o-mini , clause strength
0.254
1.1×10−23
gpt-5 , clause strength
0.520
2.1×10−108
Gemini 3 Pro, clause strength
0.599
1.1×10−150
Appendix
Table 6: Trend tests, Kendall’s τb between configured strength and realized filler-plus-repetition rate.