Paper ID: 2209.11433

The Kriston AI System for the VoxCeleb Speaker Recognition Challenge 2022

Qutang Cai, Guoqiang Hong, Zhijian Ye, Ximin Li, Haizhou Li

This technical report describes our system for track 1, 2 and 4 of the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22). By combining several ResNet variants, our submission for track 1 attained a minDCF of 0:090 with EER 1:401%. By further incorporating three fine-tuned pre-trained models, our submission for track 2 achieved a minDCF of 0:072 with EER 1:119%. For track 4, our system consisted of voice activity detection (VAD), speaker embedding extraction, agglomerative hierarchical clustering (AHC) followed by a re-clustering step based on a Bayesian hidden Markov model and overlapped speech detection and handling. Our submission for track 4 achieved a diarisation error rate (DER) of 4.86%. The submissions all ranked the 2nd places for the corresponding tracks.

Submitted: Sep 23, 2022

Topics

Voice Activity Detection
Speech Detection
VoxCeleb Speaker Recognition Challenge

Links

arXiv PDF