eess.ASOct 6, 2026

HINTT Submission to the 2nd MLC-SLM Challenge: Comparing Cascaded and Unified Approaches to Diarization and ASR

Authors: Takanori Ashihara, Kohei Matsuura, Masato Mimura

Organizations: NTT, Inc., Japan

Abstract

This paper presents the HINTT system submitted to the 2nd Challenge and Workshop on Multilingual Conversational Speech Language Model (MLC-SLM). We address multilingual speaker-attributed ASR, where systems must determine who spoke when and what was spoken. We investigate two modeling strategies for this problem: a cascaded pipeline that combines speaker diarization with speech-LLM-based ASR, and a unified speech LLM that directly generates speaker labels, timestamps, and transcriptions. Our final submission is based on the cascaded pipeline, consisting of a fine-tuned DiariZen diarization model, a fine-tuned Qwen3-ASR model, and LLM-based generative error correction. For comparison, we also fine-tune VibeVoice-ASR as a unified model using the same official training data. All task-specific fine-tuning and model selection are performed using only the official MLC-SLM data, without external data or pseudo-labels. Experimental results demonstrate that the cascaded system remains more reliable under the MLC-SLM Task 1 conditions, while unified speech LLMs offer a promising direction for future speaker-attributed ASR.

Explore similar work

CardsList
  1. An Omnilingual-ASR-Based Speech-LLM System for the 2nd MLC-SLM Challenge

    Jul 14, 2026Shuming Fang, Shuifei ZengSpeaker DiarizationWav2Vec

  2. The Second MLC-SLM Challenge: Multilingual Conversational Speech Diarization, Recognition, and Understanding

    Sep 23, 2026Bingshen Mu, Mingchen Shao, Zhennan Lin +7Multilingual Automatic Speech RecognitionSpeaker Diarization

  3. Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech

    Jul 9, 2026Hao Wu, RongQi Han, Zhen Wang +2Multilingual Automatic Speech RecognitionSpeaker Diarization