A dubbed character keeps its identity while the voice actor changes, so character identity and speaker identity are distinct properties of one recording, yet speaker verification measures only the latter. To address this gap, we propose KyaraBench, a benchmark that scores character voice directly instead of speaker voice, built from 85 human-audited identities in a dubbed anime corpus. It poses two challenging conditions: one that swaps the performer under a fixed character, and one that fixes the performer under changing characters. Listening studies with 78 participants provide human reference scores for cross-performer verification and same-actor discrimination. Speaker-verification baselines show increased errors under these character-specific conditions. We then train KyaraEmbed, a compact encoder using multilingual character supervision, same-actor negatives, and a language-alignment term. The model achieves the best performance on all character-specific conditions in the main comparison. We release the benchmark, protocol, and encoder publicly.
Figures & tables
Figure 1: Two cases of the speaker–character confusion. A dub changes the performer under a fixed character (left: answer same ), while one actor voices different characters in two shows (right: answer different ).
Stage
Identities
Utts
Clips
Anim-400K aligned dub pairs
—
436,887
858,986
show-level clustering
9,399
—
—
matched to a character portrait
4,907
—
—
≥ 3 agreeing votes, cast-listed
385
—
—
coverage and quality gates
114
1,828
3,656
human annotation (released)
85
986
1,616
Table 1: Construction funnel from Anim-400K to the released benchmark. Utts counts Japanese source utterances; Clips counts retained Japanese and English audio files.
Condition
Confusion tested
Trials (+/ − )
Mono
same char, same language
1778 / 5327
XLing
same char, dubbed performer
646 / 1812
Mono-VA
diff char, shared actor
— / 309
XLing-K
bundled, unconstrained impostor
636 / 1908
XLing-KP
bundled, pitch-matched impostor
648 / 1944
retrieval
Japanese query → English gallery
424 queries
Table 2: Conditions of KyaraBench (positives / impostors). Mono pools the Japanese and English monolingual trials. Mono-VA adds impostors only and is scored against the Mono positives. Retrieval is a 53 -way gallery (chance R@1 =.019 ).
System
Mono
XLing
Mono-VA
XLing-K
XLing-KP
R@1
shortcut
F0 statistics
.395
.418
.392
.338
.477
.087
zero-shot
ECAPA-TDNN (VoxCeleb)
.138
.367
.305
.344
.437
.130
ReDimNet-b2 (VoxCeleb2)
.156
.383
.295
.316
.418
.068
WavLM-base+ SV
.201
.314
.288
.267
.420
.052
Qwen3-Omni audio tower
.387
.455
.417
.356
.465
.137
domain
ECAPA-TDNN (domain fine-tuned)
.149
.271
.252
.225
.344
.210
Table 3: KyaraBench results (EER; lower is better, except R@1). Conditions as in Table 2 . † : paired identity-bootstrap 95% interval clear of zero against every speaker-verification baseline (ECAPA, ReDimNet, WavLM-SV). Best in bold.
BER ↓
impostor rejection ↑
Mono
XLing
K4
same-VA
diff-VA
penalty
Human
.226
.318
.319
.820
.907
+.087
ECAPA-TDNN
.143
.375
.375
.620
.820
+.200
domain FT
.107
.232
.292
.720
.820
+.100
KyaraEmbed
.143
.214
.229
.760
.780
+.020
Table 4: Human reference on identical items, from two listening studies. Left: balanced error rate (BER) on 320 verification trials (48 listeners). Right: the share of impostor trials correctly rejected in the same-actor study (200 trials, 30 listeners), for impostors that share the voice actor (same-VA) and for pitch-matched impostors with a different actor (diff-VA); the penalty is the drop in that rejection rate when the actor is shared (diff-VA - same-VA).
Configuration
Mono
XLing
Mono-VA
XLing-K
XLing-KP
R@1
domain fine-tuned
.149
.271
.252
.225
.344
.210
+ internal-dataset SupCon (adapted)
.196
.258
.281
.190
.323
.226
+ anime corpus, plain SupCon
.154
.234
.249
.165
.287
.321
+ actor-tied sampling, no margins
.134
.266
.239
.186
.307
.262
+ margins, untied characters
.133
.263
.246
.181
.306
.274
+ tie and margins ( λ=0 )
.136
.269
.243
.186
.303
.264
Table 5: Ablation ladder from the domain fine-tuned model to KyaraEmbed on KyaraBench (conditions as in Table 2 ; best in bold). The first two rungs change the training data; every rung below the second rule keeps the full corpus, the initialisation from the plain-SupCon rung, and the 1,500-step schedule, toggling one ingredient at a time.
Figure 2: Mono vs. XLing EER for 31 fine-tuned configurations from our ablations (grey), the domain fine-tuned model, the speaker-verification baselines, and our model. The dashed line shows the empirical Pareto frontier among the configurations plotted.
As audio-first agents become increasingly common in physical AI, conversational robots, and screenless wearables, audio large language models (audio-LLMs) must integrate speaker-specific understanding to support user authorization, personalization, and context-aware interaction. This requires modeling who is speaking, how the voice sounds, and how recording conditions affect speaker cues. Conventional speaker verification systems provide strong scalar scores but little linguistic evidence, while current audio-LLMs and speaker-aware language models have limited ability to organize speaker information beyond binary labels or descriptive profiles. We present SpeakerLLM, a speaker-specialized audio-LLM framework that unifies single-utterance speaker profiling, recording-condition understanding, utterance-pair speaker comparison, and evidence-organized verification reasoning within a natural-language interface. We construct verification-reasoning targets and a decision-composition policy that separate profile-level evidence from the final same-or-different decision and organize recording condition, profile evidence, and the decision into a structured trace. At its core, SpeakerLLM uses a hierarchical speaker tokenizer designed to capture multiple granularities of speaker evidence. Utterance-level speaker embeddings summarize identity and profile-level cues, whereas frame-level speaker features preserve fine-grained acoustic descriptors. Experiments show that SpeakerLLM-Base improves speaker-profile and recording-condition understanding over general audio-LLMs, while SpeakerLLM-VR preserves strong generated-verdict accuracy and produces decision traces grounded in the supervised verification reasoning schema. We will release the metadata-enriched supervision dataset and target-construction code for reproducibility.
KiHyun Nam, Jungwoo Heo, Siu Bae +2
Korea Advanced Institute of Science and Technology (KAIST) · University of Seoul
Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex storyline often relies on \textbf{speaker recognition}, the task of accurately attributing each spoken utterance to its respective character. In this paper, we advance this field through two primary contributions. (1) We introduce \textbf{DramaSR-532K}, a large-scale benchmark comprising 532K annotated dialogue lines across more than 900 unique characters, necessitating the integration of auditory, linguistic, and visual cues for speaker recognition. (2) We propose \textbf{DramaSR-LRM}, a robust approach built upon a large reasoning model (LRM). DramaSR-LRM is designed to autonomously aggregate contextual evidence via multimodal tool-use, synthesizing diverse inputs to achieve high-fidelity attribution. Experimental results demonstrate that DramaSR-LRM significantly outperforms existing baselines, particularly on short utterances where acoustic biometrics are inherently unreliable. \textit{All the data and code will be made publicly available at the project page: https://www.github.com/198808xc/DramaSR-LRM.}
Yuxuan Li, Lingxi Xie, Xinyue Huo +6
Tsinghua University, China · Huawei Inc., China · University of Chinese Academy of Sciences, China +1
While Omni-modal Large Language Models have made strides in joint sensory processing, they fundamentally struggle with a cornerstone of human interaction: deciphering complex, multi-person conversational dynamics to accurately answer Who said what and when.'' Current models suffer from an illusion of competence'' -- they exploit visual biases in conventional benchmarks to bypass genuine cross-modal alignment, while relying on sparse, low-frame-rate visual sampling that destroys crucial high-frequency dynamics like lip movements. To shatter this illusion, we introduce Visual-Registered Speaker Diarization and Recognition (VR-SDR) and the HumanOmni-Speaker Benchmark. By strictly eliminating visual shortcuts, this rigorous paradigm demands true end-to-end spatio-temporal identity binding using only natural language queries. To overcome the underlying architectural perception gap, we propose HumanOmni-Speaker, powered by a Visual Delta Encoder. By sampling raw video at 25 fps and explicitly compressing inter-frame motion residuals into just 6 tokens per frame, it captures fine-grained visemes and speaker trajectories without triggering a catastrophic token explosion. Ultimately, HumanOmni-Speaker demonstrates strong multimodal synergy, natively enabling end-to-end lip-reading and high-precision spatial localization without intrusive cropping, and achieving superior performance across a wide spectrum of speaker-centric tasks.
Detao Bai, Zhiheng Ma, Xihan Wei
Tongyi Lab Alibaba Group · Shenzhen University of Advanced Technology · Guangdong Provincial Key Laboratory of Computility Microelectronics +1