Multi-user voice agents must track who said what across dialogue sessions. Text LLMs are attractive backbones for such agents, but transcripts alone do not expose acoustic speaker identity, leaving the model without a persistent reference for linking information to speakers across sessions. We address this gap by introducing Speaker Handles, soft-token representations that expose acoustic speaker identity to a frozen text LLM for cross-session speaker-dependent reasoning. A three-stage curriculum trains a lightweight projector, with fewer than 0.1% of the backbone's parameters, to map speaker embeddings into these handles. Establishing whether the resulting handles truly support cross-session speaker-dependent reasoning is challenging with existing benchmarks because textual cues can partially reveal fact ownership. We therefore present SpeakerBind, a controlled shared-agent benchmark in which overlapping facts across users require correct cross-session speaker attribution. Speaker Handles achieve 97.40-98.36% accuracy on VoxCeleb1 and 70.40% on SpeakerBind, close to the 71.88% topline. These results show that the proposed Speaker Handles provide an efficient way to integrate acoustic speaker identity into frozen text LLMs for speaker-content reasoning.
Figures & tables
Figure 1: Speaker Handles as reusable speaker references for multi-speaker reasoning with a frozen text LLM.
Illustrative input
Turn 1 [ HA(1) ]: “The train was late.”
Turn 2 [ HB(1) ]: “We waited outside.”
Turn 3 [ HA(2) ]: “It arrived after noon.”
Optional roster [R]: Ava =HA(r) , Ben =HB(r) .
Task family
Representative training task
Stage 1: Handle grounding
Table 1: Training inputs and representative tasks.
Figure 2: Conceptual examples of SpeakerBind, where each task requires reasoning over speaker–content bindings.
VoxCeleb1
M3-SLU
FANToM
NSF-QA
SpeakerBind
Settings
O
E
H
Acc.
Fact QA
ToM
QA
Sum. (/5)
Agg.
Temp.
Follow-up
Overall
Ground-truth speaker labels (topline)
100.00
100.00
100.00
65.25
85.06
65.71
84.39
3.04
56.90
80.05
78.70
71.88
Speaker Handles (Ours)
98.36
98.36
97.40
66.43
78.51
63.34
83.69
2.83
52.60
80.00
78.60
70.40
Cascaded speaker identification
–
–
–
60.05
71.38
57.18
84.16
2.78
55.50
77.95
76.40
69.95
Transcript only (baseline)
–
–
–
53.41
52.18
47.39
81.76
2.28
0.20
25.43
47.00
24.21
Reassigned speaker bindings †
–
–
–
51.41
42.41
48.85
66.80
1.90
0.00
0.25
8.60
2.95
Table 2: Results on existing benchmarks and SpeakerBind. The three interventions disrupt speaker structure in complementary ways: Reassigned ( † ) changes the speaker assignments of answer-relevant historical turns while leaving the query speaker fixed; Same ( ‡ ) maps all turns to one shared handle; Distinct ( § ) assigns a unique handle to every turn, eliminating cross-turn identity recurrence. For SpeakerBind, expected original-gold accuracy under reassignment is 0.00, while task-specific chance levels for Agg./Temp./Follow-up are 0.00/37.05/50.00.
SpeakerBind
NSF-QA
Condition
Vox1-H
Agg.
Temp.
Follow-up
QA
Sum.
Stage 1 only
93.14
47.20
73.58
79.00
84.54
2.79
Stage 1 + 2
97.24
52.40
75.15
79.80
82.98
2.83
All at once
91.09
7.00
40.85
66.20
56.25
2.67
Full (Ours)
97.40
52.60
80.00
78.60
83.69
2.83
Table 3: Ablation of the staged training curriculum across speaker verification and downstream reasoning tasks.
As audio-first agents become increasingly common in physical AI, conversational robots, and screenless wearables, audio large language models (audio-LLMs) must integrate speaker-specific understanding to support user authorization, personalization, and context-aware interaction. This requires modeling who is speaking, how the voice sounds, and how recording conditions affect speaker cues. Conventional speaker verification systems provide strong scalar scores but little linguistic evidence, while current audio-LLMs and speaker-aware language models have limited ability to organize speaker information beyond binary labels or descriptive profiles. We present SpeakerLLM, a speaker-specialized audio-LLM framework that unifies single-utterance speaker profiling, recording-condition understanding, utterance-pair speaker comparison, and evidence-organized verification reasoning within a natural-language interface. We construct verification-reasoning targets and a decision-composition policy that separate profile-level evidence from the final same-or-different decision and organize recording condition, profile evidence, and the decision into a structured trace. At its core, SpeakerLLM uses a hierarchical speaker tokenizer designed to capture multiple granularities of speaker evidence. Utterance-level speaker embeddings summarize identity and profile-level cues, whereas frame-level speaker features preserve fine-grained acoustic descriptors. Experiments show that SpeakerLLM-Base improves speaker-profile and recording-condition understanding over general audio-LLMs, while SpeakerLLM-VR preserves strong generated-verdict accuracy and produces decision traces grounded in the supervised verification reasoning schema. We will release the metadata-enriched supervision dataset and target-construction code for reproducibility.
KiHyun Nam, Jungwoo Heo, Siu Bae +2
Korea Advanced Institute of Science and Technology (KAIST) · University of Seoul
While Omni-modal Large Language Models have made strides in joint sensory processing, they fundamentally struggle with a cornerstone of human interaction: deciphering complex, multi-person conversational dynamics to accurately answer Who said what and when.'' Current models suffer from an illusion of competence'' -- they exploit visual biases in conventional benchmarks to bypass genuine cross-modal alignment, while relying on sparse, low-frame-rate visual sampling that destroys crucial high-frequency dynamics like lip movements. To shatter this illusion, we introduce Visual-Registered Speaker Diarization and Recognition (VR-SDR) and the HumanOmni-Speaker Benchmark. By strictly eliminating visual shortcuts, this rigorous paradigm demands true end-to-end spatio-temporal identity binding using only natural language queries. To overcome the underlying architectural perception gap, we propose HumanOmni-Speaker, powered by a Visual Delta Encoder. By sampling raw video at 25 fps and explicitly compressing inter-frame motion residuals into just 6 tokens per frame, it captures fine-grained visemes and speaker trajectories without triggering a catastrophic token explosion. Ultimately, HumanOmni-Speaker demonstrates strong multimodal synergy, natively enabling end-to-end lip-reading and high-precision spatial localization without intrusive cropping, and achieving superior performance across a wide spectrum of speaker-centric tasks.
Detao Bai, Zhiheng Ma, Xihan Wei
Tongyi Lab Alibaba Group · Shenzhen University of Advanced Technology · Guangdong Provincial Key Laboratory of Computility Microelectronics +1
Voice provides a natural and immediate interface for AI agents. Many settings in which voice agents could be useful, including meetings, households, and collaborative work, are inherently multi-speaker. Supporting these settings introduces challenges that are largely absent from one-on-one interaction. We introduce the Multi-Speaker Interaction Benchmark (MSI-Bench) for evaluating multi-speaker voice interaction. Each test case is a short multi-party multi-turn audio scene with participant context, expected tool calls, and atomic rubrics. The benchmark targets three capability families: multi-speaker memory, multi-speaker instruction following, and multi-speaker reasoning. It comprises 1,152 test cases, evenly split between Mandarin Chinese and English (576 each). The strongest configuration on each split passes all rubrics on only 66.8% of English and 54.5% of Mandarin cases, and the strongest open-weight configuration on 34.0% and 19.3%. Failure analysis separates perception from reasoning: open-weight models are bottlenecked by the multi-speaker audio front-end, while frontier systems still fail speaker-scoped decision making on clean transcripts---and models across the board often respond when no one has addressed them. These results identify speaker-grounded perception, speaker-scoped decision making, and conversational restraint as concrete targets for future voice agents.