eess.ASOct 5, 2026

Character Identity is not Speaker Identity: KyaraBench and KyaraEmbed for Character Verification

Authors: Joonyong Park, Jerry Li

Organizations: Spellbrush, CA, USA

Abstract

A dubbed character keeps its identity while the voice actor changes, so character identity and speaker identity are distinct properties of one recording, yet speaker verification measures only the latter. To address this gap, we propose KyaraBench, a benchmark that scores character voice directly instead of speaker voice, built from 85 human-audited identities in a dubbed anime corpus. It poses two challenging conditions: one that swaps the performer under a fixed character, and one that fixes the performer under changing characters. Listening studies with 78 participants provide human reference scores for cross-performer verification and same-actor discrimination. Speaker-verification baselines show increased errors under these character-specific conditions. We then train KyaraEmbed, a compact encoder using multilingual character supervision, same-actor negatives, and a language-alignment term. The model achieves the best performance on all character-specific conditions in the main comparison. We release the benchmark, protocol, and encoder publicly.

Figures & tables

Explore similar work

CardsList
  1. SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning

    May 14, 2026KiHyun Nam, Jungwoo Heo, Siu Bae +2SpeakerLarge Audio Language Models

  2. Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

    Jul 2, 2026Yuxuan Li, Lingxi Xie, Xinyue Huo +6SpeakerVideo Understanding

  3. HumanOmni-Speaker: Identifying Who said What and When

    Mar 23, 2026Detao Bai, Zhiheng Ma, Xihan WeiSpeaker DiarizationConversational Context