Towards an Extensible Benchmark for Spoken Dialogue with Social Robots
Authors: Casey Kennington, Ross Mead, Saad Elbeleidy, Jesse Thomason
Organizations: Department of Computer Science Boise State University Boise, Idaho, U.S.A. · Semio Los Angeles, California, U.S.A. · Peerbots Arlington, Virginia, U.S.A. · School of Interactive Computing Georgia Institute of Technology Atlanta, Georgia, U.S.A.
Language models provide a plug-and-play interface between humans and robots, but important challenges remain when speech, dialogue, fast interaction, and collaboration are required. We propose a benchmark for the community to use as a way to explore common spoken dialogue artifacts between robots and humans, including requests for clarification, interruptions, embodied signals (e.g., head nods or facial cues), and time constraints. We also explain our vision to extend the benchmark for other aspects of human-robot interaction that are important to the larger research community. To facilitate the benchmark, we further propose using \textit{Retico}, a real-time communication framework that fulfills important technical requirements to enable robots to have spoken dialogue capabilities.
Figures & tables
Robot-Ready SDS
LLM-Based Interaction
speech-driven
text-driven
word-level granularity
sentence-level granularity
system requests clarification
system assumes understanding
turn-taking is fluid
text-driven turns
live feedback
feedback not possible
multimodal inputs are continuous
multimodal inputs are tokenized
TABLE I: Technical distinctions between robot-ready SDS and LLM-based interaction.
Fig. 1: A top-down view of the base task. The game board is on the bottom which is visible to the instruction follower. Behind a divider is the target shape, the robot, and on either side of the robot are two puzzle pieces, at least one of them required for completing the task.
Fig. 2: The human / instruction follower view of the game. The human can see the robot’s face and arms, as well as the game board, but not the target shape or the two hidden pieces. The human can easily reach around the divider to replace one of the two hidden pieces.
Graduate School of Data Science, Seoul National University · Department of Information Systems, Hanyang University · Department of Computer Science and Engineering, Seoul National University