Clinical research in psychiatry increasingly relies on large scale collection of spoken language data to identify acoustic and linguistic biomarkers. Yet evolving consent and protocol requirements can oblige investigators to remove a designated speaker from multi-speaker recordings and to verify said removal at a scale infeasible for manual review of entire corpora. We study this verification problem for role-driven dyadic clinical dialogue in psychiatry and investigate it with two parallel, symmetric pipelines: confirming that clinician speech has been removed from psychiatric interview recordings, and confirming that patient speech has been removed from the same recordings. Each pipeline redacts the raw audio for its target role and then scans the surviving output with audio-language and large-language models to identify missed deletions. We evaluate this approach on a corpus of 48 dyadic recordings drawn from psychiatry settings, testing four open-weight models in an inference-only setting: Gemma-4-12B, Gemma-4-31B, Nemotron-3-Nano, and Nemotron-3-Nano-Omni. A disjunctive OR ensemble over fourteen model-view configurations had a combined F1 of 0.478 (precision 0.330, recall 0.870), an improvement over individual model estimates driven by recall gains that point to substantial complementarity across models and context views.
Figures & tables
Metric
Clinician
Patient
Sessions with no missed deletions
9
5
Total annotated spans
515
872
Spans / session (mean)
10.73
18.17
Spans / session (median)
5.5
13.5
Spans / session (range)
0–50
0–54
Annotated seconds / session (mean)
23.4
33.1
Table 1: Redaction annotation summary across 48 sessions.
Patient Del.
Clinician Del.
Combined
Model (mode)
P
R
F1
P
R
F1
P
R
F1
Audio only (20s)
Gemma-4-12B
0.057
0.302
0.096
0.038
0.153
0.061
0.051
0.247
0.085
Nemotron-3-Nano-Omni
0.050
0.495
0.090
0.031
0.311
0.056
0.043
0.427
0.077
Audio-text pairs (20s)
Gemma-4-12B
0.273
0.718
0.396
0.242
0.511
0.329
0.263
0.641
0.373
Table 2: Individual model performance, micro-averaged over 48 sessions, broken out by patient deletion, clinician deletion, and combined (pooled) pipelines. For each view, the P/R/F1 triple with the highest F1 is bolded.
Patient Del.
Clinician Del.
Combined
Ensemble
P
R
F1
P
R
F1
P
R
F1
OR (any)
0.277
0.943
0.428
0.138
0.864
0.238
0.205
0.913
0.334
OR (no-audio)
0.368
0.868
0.517
0.257
0.773
0.385
0.320
0.833
0.462
OR (with-audio)
0.208
0.845
0.334
0.098
0.672
0.171
0.153
0.781
0.256
OR (transcript-guided)
0.380
0.898
0.534
0.265
0.823
0.401
0.330
0.870
0.478
OR (audio-text pairs)
0.305
0.772
0.437
0.238
0.610
0.342
0.280
0.712
0.402
Table 3: Ensemble performance, micro-averaged over 48 sessions, broken out by patient deletion, clinician deletion, and combined (pooled) pipelines. Ensembles combine the configurations in Table 2 via disjunction (OR) or majority vote. The P/R/F1 triple with the highest F1 is bolded for each view.
Automated depression screening from clinical interviews requires attribution of utterances to the clinician or patient. We evaluate two datasets: DAIC-WOZ, where participant-only recordings require re-synthesizing both sides for controlled two-party evaluation, and PDCH-HAMD, comprising voice-converted real Chinese interviews for cross-lingual validation. Cascaded systems combine speaker diarization with role-assignment heuristics, so errors can propagate across stages. We propose an end-to-end model, which we named DiaWhisper, that fine-tunes Whisper-large-v3 with LoRA and an auxiliary frame-level role head for transcription and attribution, together with DiaWhisper-DPO, a failure-mined refinement that uses genuine decoding failures as DPO rejected completions without human preference annotation. On 29 DAIC-WOZ test sessions, DiaWhisper-DPO achieves 0.973 role accuracy and 0.119 DER, 72% below the strongest cascaded baseline, and reduces seed variation from σ = .205 to .002. Retrained on PDCH-HAMD, it achieves 0.757 role accuracy and improves all 78 session-seed pairs.
Weiming Li, Ana Catarina Fidalgo Barata, Miguel Constante +1
Institute for Systems and Robotics (ISR), LARSyS, Departamento de Bioengenharia, Instituto Superior Técnico (IST), Universidade de Lisboa, 1049-001 Lisboa, Portugal · Hospital Beatriz Ângelo, Faculdade de Medicina Universidade Católica Portuguesa, 2674-514 Loures, Portugal
As audio-first agents become increasingly common in physical AI, conversational robots, and screenless wearables, audio large language models (audio-LLMs) must integrate speaker-specific understanding to support user authorization, personalization, and context-aware interaction. This requires modeling who is speaking, how the voice sounds, and how recording conditions affect speaker cues. Conventional speaker verification systems provide strong scalar scores but little linguistic evidence, while current audio-LLMs and speaker-aware language models have limited ability to organize speaker information beyond binary labels or descriptive profiles. We present SpeakerLLM, a speaker-specialized audio-LLM framework that unifies single-utterance speaker profiling, recording-condition understanding, utterance-pair speaker comparison, and evidence-organized verification reasoning within a natural-language interface. We construct verification-reasoning targets and a decision-composition policy that separate profile-level evidence from the final same-or-different decision and organize recording condition, profile evidence, and the decision into a structured trace. At its core, SpeakerLLM uses a hierarchical speaker tokenizer designed to capture multiple granularities of speaker evidence. Utterance-level speaker embeddings summarize identity and profile-level cues, whereas frame-level speaker features preserve fine-grained acoustic descriptors. Experiments show that SpeakerLLM-Base improves speaker-profile and recording-condition understanding over general audio-LLMs, while SpeakerLLM-VR preserves strong generated-verdict accuracy and produces decision traces grounded in the supervised verification reasoning schema. We will release the metadata-enriched supervision dataset and target-construction code for reproducibility.
KiHyun Nam, Jungwoo Heo, Siu Bae +2
Korea Advanced Institute of Science and Technology (KAIST) · University of Seoul
Automatic depression detection from doctor-patient conversations has gained momentum thanks to the availability of public corpora and advances in language modeling. However, interpretability remains limited: strong performance is often reported without revealing what drives predictions. We analyze three datasets: ANDROIDS, DAIC-WOZ, E-DAIC and identify a systematic bias from interviewer prompts in semi-structured interviews. Models trained on interviewer turns exploit fixed prompts and positions to distinguish depressed from control subjects, often achieving high classification scores without using participant language. Restricting models to participant utterances distributes decision evidence more broadly and reflects genuine linguistic cues. While semi-structured protocols ensure consistency, including interviewer prompts inflates performance by leveraging script artifacts. Our results highlight a cross-dataset, architecture-agnostic bias and emphasize the need for analyses that localize decision evidence by time and speaker to ensure models learn from participants' language.
Hasindri Watawana, Sergio Burdisso, Diego A. Moreno-Galván +4
Idiap Research Institute, Switzerland · EPFL, Switzerland · Centro de Investigación en Matemáticas (CIMAT), Mexico +1