cs.CLOct 6, 2026

Language Carries the Expert's Impression: Instrument-Anchored LLM Judges Transfer Counseling-Quality Assessment and Beat In-Domain Training

Authors: Tobias Hallmen, Elisabeth André

Organizations: Chair for Human-Centered Artificial Intelligence University of Augsburg

Abstract

Automatic assessment of communication quality in dyadic counseling conversations is bottlenecked by data: expert-rated corpora are small and expensive to grow. We study cross-domain transfer of expert overall-impression prediction across three German corpora of simulated counseling (two general-practice medical, one school-related parent-teacher; n=195n=195 expert-rated sessions, one corpus after scale equating). Training on the other domains beats training in-domain: leave-one-domain-out transfer reaches nested Spearman ρ=0.54ρ= 0.54 against ≤0.48\le 0.48 within the target domain, a paired session-level gap of +0.15+0.15 that holds at +0.12+0.12 when the training-set sizes are matched, so it is not simply data volume. The decisive features are session-level construct scores from small open-weight LLMs reading the two-speaker transcript, with the constructs largely derived from the experts' rating instruments: the instrument-derived battery lifts a single judge from 0.320.32 to 0.410.41 over generic dialogue qualities, judges from three model families ensemble to 0.510.51 language-only, and a nonverbal-dyadic block adds +0.03+0.03 more, not separable from noise at this sample size. We also price the recording setup: one corpus lost its per-speaker audio, 16% of its diarised segments carry the wrong speaker, and repair is worth +0.07+0.07 there. At practically attainable corpus sizes, the expert's overall impression is carried by what is said, and by other communication programs' data more than by one's own.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Structured Prompting and Automated Evaluation in Fixed Synthetic Japanese-Language Counseling Dialogues

    Jun 28, 2025Keita Kiuchi, Yoshikazu Fujimoto, Hideyuki Goto +5CounselingPrompt Engineering

  2. SpeechCritic: Learning a Diagnostic Speech Judge from Limited Human Preferences

    Sep 28, 2026Mingyue Huo, Shivam Mehta, Bhavin Jawade +2Large Audio Language ModelsCritic

  3. A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

    Jul 8, 2026A. Sayyad, J. Emmons, S. Jones +2RaterFull-Duplex Voice Agents