eess.ASSep 29, 2026

Voices as Handles: Reasoning about Speaker Identity with Frozen Text LLMs

Authors: Runqiu Xu, Zhisheng Zheng, David Harwath

Organizations: The University of Texas at Austin, Austin, TX, USA

Abstract

Multi-user voice agents must track who said what across dialogue sessions. Text LLMs are attractive backbones for such agents, but transcripts alone do not expose acoustic speaker identity, leaving the model without a persistent reference for linking information to speakers across sessions. We address this gap by introducing Speaker Handles, soft-token representations that expose acoustic speaker identity to a frozen text LLM for cross-session speaker-dependent reasoning. A three-stage curriculum trains a lightweight projector, with fewer than 0.1% of the backbone's parameters, to map speaker embeddings into these handles. Establishing whether the resulting handles truly support cross-session speaker-dependent reasoning is challenging with existing benchmarks because textual cues can partially reveal fact ownership. We therefore present SpeakerBind, a controlled shared-agent benchmark in which overlapping facts across users require correct cross-session speaker attribution. Speaker Handles achieve 97.40-98.36% accuracy on VoxCeleb1 and 70.40% on SpeakerBind, close to the 71.88% topline. These results show that the proposed Speaker Handles provide an efficient way to integrate acoustic speaker identity into frozen text LLMs for speaker-content reasoning.

Figures & tables

Explore similar work

CardsList
  1. SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning

    May 14, 2026KiHyun Nam, Jungwoo Heo, Siu Bae +2SpeakerLarge Audio Language Models

  2. HumanOmni-Speaker: Identifying Who said What and When

    Mar 23, 2026Detao Bai, Zhiheng Ma, Xihan WeiSpeaker DiarizationConversational Context

  3. MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents

    Sep 21, 2026Chenxu Xiong, Dongming Shen, Yuzhi Tang +3Full-Duplex Voice AgentsDialogue Benchmarks