cs.SDSep 30, 2026

Training-Free Affinity Fusion of Neural and Embedding-Based Speaker Diarization

Authors: Yehoshua Dissen, Joseph Keshet, Eduard Golshtein

Organizations: Linguana · Technion – Israel Institute of Technology

Abstract

Speaker diarization systems based on speaker embeddings and neural diarization exploit complementary forms of speaker information, but their intermediate representations are not directly compatible. We introduce Training-Free Affinity Fusion (TFAF), which integrates the speaker structure inferred by a neural diarizer into an embedding-based diarization system. The neural speaker partition is used to condition local speaker representations, from which we construct a continuous affinity matrix and combine it with the embedding-based acoustic affinity before a single global clustering step. The method requires no additional training, shared embedding space, speaker-label alignment, or hard transfer of the neural diarizer's speaker count. Experiments on AMI and CALLHOME show consistent DER improvements over both constituent systems; on AMI, fusion also improves speaker-attributed transcription. Ablations show that the neural speaker partition accounts for most of the gain, while retaining the continuous embedding-based affinities provides additional benefit over hard partition fusion.

Figures & tables

Explore similar work

CardsList
  1. DiariZen Explained: A Tutorial for the Open Source State-of-the-Art Speaker Diarization Pipeline

    Apr 23, 2026Nikhil RaghavSpeaker DiarizationSpeaker

  2. TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding

    Jan 11, 2026Mingyue Huo, Yiwen Shao, Yuheng ZhangSpeaker DiarizationSpeech-To-Text Alignment