cs.CLOct 7, 2026

Prompts versus Rules: Auditing and Controlling Speech Naturalness Behaviors in Voice User Simulators

Authors: Riqiang Wang, Elena Khasanova, Harsh Saini, Lex Konnelly, Parsa Kavehzadeh, Matthias Lee, Mohamed Attia

Organizations: Dialpad Inc.

Abstract

As voice agents gain more popularity commercially, the user simulators used to evaluate the deployed agents are also being developed to include more realistic, variable, and diverse speech naturalness behaviors -- disfluency, interruption and backchanneling. The quality of the user simulator directly affects the validity of agent evaluation results. However, we find that most studies so far have not examined in detail whether the intended configuration for these behaviors is realized in the simulation. In this study, we audit the realized naturalness behaviors of tau-Voice, our own LLM-based prompting approach across three models, and our rule-based injection algorithm for disfluency, interruption, and backchanneling. We find that prompting for these behaviors is unreliable and produces speech inconsistent with the instructions, placed and distributed less naturally than the instruction implies. In contrast, our rule-based, model-free algorithm produces controllable and diverse naturalness behaviors more aligned with natural speech. Our results suggest that LLMs not purpose-trained for user simulation are not sufficient on their own to represent authentic user behavior, and that linguistically informed deterministic approaches or specialized models are needed to close the gap; auditing and reporting realized naturalness behaviors, rather than configured settings, is what makes that gap visible.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Investigating Assistant Bias in LLM User Simulators Using a Role Vector

    Sep 1, 2026Daeheon Jeong, Yoonjoo Lee, Eugene Choi +2Large Language Model-Based Role-Play SimulationUser Simulation

  2. Controllable User Simulation

    May 12, 2026Guy Tennenholtz, Ofer Meshi, Amir Globerson +3Off-Policy EvaluationCausal Inference

  3. Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation

    Aug 10, 2026Bo Wang, Ruixing Zhang, Yunqi Liu +4User SimulationConditional Generation