cs.LGSep 29, 2026

Role-guided Speaker Deletion Verification in Clinical Psychiatry Speech Recordings with Audio Language Models

Authors: Joseph T Colonel, Daniel Katzman, Kelsey Kirker, Adam N Davidson, Shalaila S Haas, Cheryl Corcoran, René S Kahn, Guillermo Checci, +1 more

Organizations: Icahn School of Medicine at Mount Sinai, Department of Psychiatry, New York, NY, USA

Abstract

Clinical research in psychiatry increasingly relies on large scale collection of spoken language data to identify acoustic and linguistic biomarkers. Yet evolving consent and protocol requirements can oblige investigators to remove a designated speaker from multi-speaker recordings and to verify said removal at a scale infeasible for manual review of entire corpora. We study this verification problem for role-driven dyadic clinical dialogue in psychiatry and investigate it with two parallel, symmetric pipelines: confirming that clinician speech has been removed from psychiatric interview recordings, and confirming that patient speech has been removed from the same recordings. Each pipeline redacts the raw audio for its target role and then scans the surviving output with audio-language and large-language models to identify missed deletions. We evaluate this approach on a corpus of 48 dyadic recordings drawn from psychiatry settings, testing four open-weight models in an inference-only setting: Gemma-4-12B, Gemma-4-31B, Nemotron-3-Nano, and Nemotron-3-Nano-Omni. A disjunctive OR ensemble over fourteen model-view configurations had a combined F1 of 0.478 (precision 0.330, recall 0.870), an improvement over individual model estimates driven by recall gains that point to substantial complementarity across models and context views.

Figures & tables

Explore similar work

CardsList
  1. DiaWhisper-DPO: Role-Attributed Transcription of Clinical Interviews via Failure-Mined Preference Optimization

    Sep 15, 2026Weiming Li, Ana Catarina Fidalgo Barata, Miguel Constante +1Speaker DiarizationTranscript

  2. SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning

    May 14, 2026KiHyun Nam, Jungwoo Heo, Siu Bae +2SpeakerLarge Audio Language Models

  3. When Consistency Becomes Bias: Interviewer Effects in Semi-Structured Clinical Interviews

    Mar 25, 2026Hasindri Watawana, Sergio Burdisso, Diego A. Moreno-Galván +4DepressionInterviews