cs.SDSep 30, 2026

Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across Meetings

Authors: Shantanu Vispute, Aditya Mishra, Siddhartha Saxena

Organizations: Foyer

Abstract

Speech transcripts used as long-term memory must preserve both words and stable speaker identities. Existing meeting-transcription metrics either ignore speakers or remap anonymous speakers independently in each recording, so they cannot measure whether the same person retains one identity across meetings. We evaluate persistent speaker attribution with Speaker Identified cpWER (SI-cpWER), which scores a corpus under one global speaker-ID assignment. The benchmark covers five commercial diarize-then-identify cascades, two open academic baselines, and ThyVoice on the full 129-meeting CHiME-8 NOTSOFAR evaluation set in clean and noiseaugmented form, plus CHiME-6. ThyVoice is our end-to-end reference system; it repairs overlap and gates the evidence used to create and update voiceprints. Requiring persistent identity changes the commercial ranking: ThyVoice records lower SI-cpWER than every evaluated commercial cascade in all three conditions and the lowest mean in the full panel, 47.13 versus 54.75 for the next system. Complementary lexical, diarization, per-recording attribution, and speaker-clustering diagnostics characterize upstream error surfaces in the final attributed record. These results show why persistent attribution must be evaluated directly in systems that reuse conversations across time.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. STAM-ASR: Speaker-Temporal Anchoring with Memory for Multi-Speaker ASR

    Sep 24, 2026Victor Tolulope Olufemi, Syeda Faiza Ahmed Sara, Shammur Absar ChowdhuryAutomatic Speech RecognitionSpeaker

  2. TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding

    Jan 11, 2026Mingyue Huo, Yiwen Shao, Yuheng ZhangSpeaker DiarizationSpeech-To-Text Alignment

  3. Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR

    Jun 28, 2026Yichi Wang, Junzhe Chen, Wangjin Zhou +1Target Speaker ExtractionSpeaker