Sep 15, 2026 · cs.CLJ/K move · Enter open · S save
Weiming Li, Ana Catarina Fidalgo Barata, Miguel Constante, João Miguel Sanches
Institute for Systems and Robotics (ISR), LARSyS, Departamento de Bioengenharia, Instituto Superior Técnico (IST), Universidade de Lisboa, 1049-001 Lisboa, Portugal · Hospital Beatriz Ângelo, Faculdade de Medicina Universidade Católica Portuguesa, 2674-514 Loures, Portugal
Automated depression screening from clinical interviews requires attribution of utterances to the clinician or patient. We evaluate two datasets: DAIC-WOZ, where participant-only recordings require re-synthesizing both sides for controlled two-party evaluation, and PDCH-HAMD, comprising voice-converted real Chinese interviews for cross-lingual validation. Cascaded systems combine speaker diarization with role-assignment heuristics, so errors can propagate across stages. We propose an end-to-end model, which we named DiaWhisper, that fine-tunes Whisper-large-v3 with LoRA and an auxiliary frame-level role head for transcription and attribution, together with DiaWhisper-DPO, a failure-mined refinement that uses genuine decoding failures as DPO rejected completions without human preference annotation. On 29 DAIC-WOZ test sessions, DiaWhisper-DPO achieves 0.973 role accuracy and 0.119 DER, 72% below the strongest cascaded baseline, and reduces seed variation from σ = .205 to .002. Retrained on PDCH-HAMD, it achieves 0.757 role accuracy and improves all 78 session-seed pairs.