Towards an Extensible Benchmark for Spoken Dialogue with Social Robots
Authors: Casey Kennington, Ross Mead, Saad Elbeleidy, Jesse Thomason
Organizations: Department of Computer Science Boise State University Boise, Idaho, U.S.A. · Semio Los Angeles, California, U.S.A. · Peerbots Arlington, Virginia, U.S.A. · School of Interactive Computing Georgia Institute of Technology Atlanta, Georgia, U.S.A.
Language models provide a plug-and-play interface between humans and robots, but important challenges remain when speech, dialogue, fast interaction, and collaboration are required. We propose a benchmark for the community to use as a way to explore common spoken dialogue artifacts between robots and humans, including requests for clarification, interruptions, embodied signals (e.g., head nods or facial cues), and time constraints. We also explain our vision to extend the benchmark for other aspects of human-robot interaction that are important to the larger research community. To facilitate the benchmark, we further propose using \textit{Retico}, a real-time communication framework that fulfills important technical requirements to enable robots to have spoken dialogue capabilities.
Figures & tables
Robot-Ready SDS
LLM-Based Interaction
speech-driven
text-driven
word-level granularity
sentence-level granularity
system requests clarification
system assumes understanding
turn-taking is fluid
text-driven turns
live feedback
feedback not possible
multimodal inputs are continuous
multimodal inputs are tokenized
TABLE I: Technical distinctions between robot-ready SDS and LLM-based interaction.
Fig. 1: A top-down view of the base task. The game board is on the bottom which is visible to the instruction follower. Behind a divider is the target shape, the robot, and on either side of the robot are two puzzle pieces, at least one of them required for completing the task.
Fig. 2: The human / instruction follower view of the game. The human can see the robot’s face and arms, as well as the game board, but not the target shape or the two hidden pieces. The human can easily reach around the divider to replace one of the two hidden pieces.
Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular interface for interacting with such systems. Speech input introduces an additional failure point: transcription errors can alter task-critical entities, constraints, or targets before the agent begins reasoning, while conventional ASR metrics do not directly measure whether the information required for successful execution has been preserved. We introduce Talk2Agent, a benchmark for evaluating how effectively voice interfaces convey human-spoken instructions to LLM-based computer-use agents. Talk2Agent builds human-spoken versions of tasks from WildClawBench and OSWorld and evaluates a range of voice interfaces, including dedicated ASR models, audio-capable LLMs, contextual biasing, and LLM-based ontology repair. Because repeatedly executing long-horizon computer-use tasks is costly and stochastic, we further propose an execution-free, task-conditioned evaluation framework that projects the original task grader onto prompt-addressable intentions and measures how much task-relevant information is retained after the voice interface. On WildClawBench, Talk2Agent's execution-free native projection provides a practical, execution-grounded measure of voice-interface quality, correlating with downstream task completion and improving Pearson correlation by 0.246 over WER/CER on 32 hours of real human speech.
Terumi Chiba, Guangzhi Sun, Zheqi Yuan +1
Department of Electronic Engineering, Tsinghua University · University of Cambridge
For noisy real-world environments such as those in open public spaces, spoken dialogue systems for both autonomous robots and avatars should be carefully designed to provide enhanced speech signals. These signals can be used either for speech recognition or, in the case of an avatar system, transmitted as clean speech to a remote operator. This work proposes an audio system that can be used for both these scenarios and was demonstrated as a proof-of-concept at the 2025 World Expo in Osaka. The first scenario is an attentive listening system with the android ERICA, and the second is a conversation support system with mobile Teleco robots, with one of them acting as an avatar for a remote operator. Both systems feature multi-party conversation and use a single multi-channel microphone array. We describe how our audio system not only enhances the speech of multiple speakers in a noisy environment, but provides a form of spatial audio which allows for more immersiveness in avatar-based conversational interactions.
Divesh Lala, Yogeeswaran Muthukumaran, Vincent Fernandes +10
1Kyoto University Graduate School of Informatics · 2Osaka University, Graduate School of Engineering Science · 3National University of Singapore +1
Robust voice agents require exposure to the full diversity of how people interact through speech. However, obtaining enough spoken interactions is prohibitively expensive. Building spoken user simulators that address this requires large-scale spoken task-oriented dialogue (TOD) data encompassing spoken user behaviors, yet existing datasets are limited in scale and domain coverage, with no systematic pipeline for augmenting them. To address this, we introduce SpokenTOD, a spoken TOD dataset of 52,390 dialogues and 1,034 hours of speech augmented with four spoken user behaviors---cross-turn slots, barge-in, disfluency, and emotional prosody---across diverse speakers and domains. Building on SpokenTOD, we present SpokenUS, a spoken user simulator grounded in TOD that decides when to speak through a dedicated turn-taking head. SpokenUS achieves comparable goal coverage to much larger models while substantially outperforming all baselines in human MOS, disclosing slot values gradually across the dialogue as humans do rather than front-loading them. Further analysis confirms that SpokenUS's spoken behaviors pose meaningful challenges to voice agents, making it a practical tool for evaluating more robust spoken dialogue systems. Our code is available at https://github.com/holi-lab/SpokenUS.
Jonggeun Lee, Junseong Pyo, Jeongmin Park +1
Graduate School of Data Science, Seoul National University · Department of Information Systems, Hanyang University · Department of Computer Science and Engineering, Seoul National University