Language Carries the Expert's Impression: Instrument-Anchored LLM Judges Transfer Counseling-Quality Assessment and Beat In-Domain Training
Organizations: Chair for Human-Centered Artificial Intelligence University of Augsburg
Abstract
Automatic assessment of communication quality in dyadic counseling conversations is bottlenecked by data: expert-rated corpora are small and expensive to grow. We study cross-domain transfer of expert overall-impression prediction across three German corpora of simulated counseling (two general-practice medical, one school-related parent-teacher; expert-rated sessions, one corpus after scale equating). Training on the other domains beats training in-domain: leave-one-domain-out transfer reaches nested Spearman against within the target domain, a paired session-level gap of that holds at when the training-set sizes are matched, so it is not simply data volume. The decisive features are session-level construct scores from small open-weight LLMs reading the two-speaker transcript, with the constructs largely derived from the experts' rating instruments: the instrument-derived battery lifts a single judge from to over generic dialogue qualities, judges from three model families ensemble to language-only, and a nonverbal-dyadic block adds more, not separable from noise at this sample size. We also price the recording setup: one corpus lost its per-speaker audio, 16% of its diarised segments carry the wrong speaker, and repair is worth there. At practically attainable corpus sizes, the expert's overall impression is carried by what is said, and by other communication programs' data more than by one's own.
Figures & tables
| Ticks | Vacc | Teach | |
| sessions | 57 | 53 | 85 |
| duration (min) | 7.1 0.6 | 6.9 0.9 | 9.5 2.3 |
| assessed role | doctor | doctor | teacher |
| interlocutor role | patient | patient | parent |
| target instrument | BGR | BGR | global item |
| rating | 1–5 | 1–5 | 1–5 / 1–4 |
| LODO | CV | ||
| Feature set | fixed | best | best |
| Language: LLM judges | |||
| final ensemble (5 judges) | .528 | .528 | .453 |
| initial ensemble (4 judges) | .450 | .465 | .416 |
| granite-30B judge | .448 | .464 | .405 |
| BGR+OSCE judge | .418 | .431 | .293 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Placement of the four-point era | ties | |
|---|---|---|
| as stored | .44 | 33.2% |
| midpoint-skipping | .47 | 30.8% |
| equal-interval | .53 | 21.0% |
| shifted below | .58 | 30.4% |
| equipercentile † | .40 | 50.3% |
| within-era reference | .51 | — |
| Battery | Model (family) | #c | ens. |
|---|---|---|---|
| generic dialogue | gemma3-12B (gemma) | 5 | |
| generic dialogue | qwen3-14B, think (qwen) | 5 | |
| BGR OSCE | gemma3-12B (gemma) | 8 | |
| anti-band wording | gemma3-12B (gemma) | 8 | |
| anti-band wording | qwen3.6-27B (qwen) | 8 | |
| anti-band wording | granite4.1-30B (granite) | 8 |
| Prompt | LODO | folds |
|---|---|---|
| generic, 5 constructs | .321 | .34/.32/.30 |
| experts’ 8 constructs, bare | .414 | .45/.36/.43 |
| quoted level descriptions | .418 | .40/.43/.42 |
| Construct (judge) | Tic | Vac | Tea | min |
| phase completeness (P) | .33 | .44 | .40 | .33 |
| coherence (B) | .33 | .35 | .38 | .33 |
| overall grade (B) | .32 | .29 | .41 | .29 |
| clarity (think) | .31 | .27 | .28 | .27 |
| structure (generic) | .26 | .31 | .27 | .26 |
| empathy (generic) | .30 | .25 | .38 | .25 |
| Held out | selected inside the fold | |
|---|---|---|
| Ticks | final stack, PCA-2, ridge | .565 |
| Vacc | final stack w/o prosody, PCA-5, ridge | .489 |
| Teach | judges nonverbal, PCA-3, SVR | .570 |
| mean | .542 |
| PCA | -score | RFE | mut. inf. | headline | |
|---|---|---|---|---|---|
| 1 | .240 | .219 | .191 | .168 | .251 |
| 2 | .269 | .248 | .223 | .204 | .552 |
| 3 | .277 | .248 | .224 | .206 | .551 |
| 5 | .278 | .263 | .242 | .224 | .551 |
| 8 | .289 | .269 | .251 | .246 | .551 |
| 12 | .285 | .267 | .270 | .253 | .551 |
| Held-out | LODO | LODO m | CV m | CV | |
|---|---|---|---|---|---|
| final stack | |||||
| Ticks | 46 | .562 | .529 .048 | .384 .052 | 0.99 |
| Vacc | 42 | .483 | .442 .046 | .306 .049 | 0.98 |
| Teach | 68 | .607 | .592 .025 | .527 .023 | 0.98 |
| mean | .551 | .521 | .405 | ||
| final judges (language only) | |||||
| foreign | own | own only | |
|---|---|---|---|
| final stack | .551 | .526 | .392 |
| final judges | .528 | .507 | .423 |
| MAE | mean-only | QWK | sd(pred) | ||
|---|---|---|---|---|---|
| Ticks | .562 | .699 | .908 | .347 | 0.43 |
| Vacc | .498 | .728 | .816 | .361 | 0.49 |
| Teach | .566 | .652 | .754 | .375 | 0.38 |
| pooled | .545 | .687 | .815 | .364 | 0.43 |
| Identifier | Meaning |
|---|---|
| llm _only | judge alone: 2 generic (gemma3-12B), 3 thinking (qwen3-14B), 4 BGR OSCE (gemma3-12B), 5 phase technique, 6 anti-band BGR OSCE (Appendix B ), 7 qwen3.6-27B, 7m its 3-seed mean, 8 phi4-14B, 9 granite4.1-30B |
| llm _z | per-domain z-scored scores of those judges, concatenated |
| lowdim_all _z | the same judges plus the compact block and the dyadic deltas |
| c / f suffix | judge scores from the AU50-repaired vaccines transcripts (Section 6 ); c = every judge of the set repaired, plus the seed-mean qwen3.6 |
| pro / prosody2 | prosody block (semitone , spread/slope, jitter/shimmer/HNR) |
| shallow_ling7 | shallow transcript statistics (7 surface counts) |
| Feature set | -score | MI | RFE | PCA |
|---|---|---|---|---|
| lowdim_all23479c_pro_z | 0.518 | 0.461 | 0.467 | 0.553 |
| lowdim_all2349c_z | 0.518 | 0.512 | 0.536 | 0.549 |
| lowdim_all23479c_z | 0.499 | 0.478 | 0.449 | 0.545 |
| lowdim_all234789c_z | 0.526 | 0.487 | 0.501 | 0.536 |
| lowdim_all23479_z | 0.493 | 0.479 | 0.500 | 0.534 |
| lowdim_all2349f_z | 0.511 | 0.512 | 0.519 | 0.529 |
| Feature set | -score | MI | RFE | PCA |
|---|---|---|---|---|
| lowdim_all234789c_z | 0.356 | 0.317 | 0.323 | 0.482 |
| lowdim_all23479c_z | 0.374 | 0.372 | 0.386 | 0.471 |
| lowdim_all23479c_pro_z | 0.383 | 0.339 | 0.356 | 0.465 |
| llm234789c_z | 0.390 | 0.390 | 0.309 | 0.456 |
| llm2349c_z | 0.371 | 0.342 | 0.321 | 0.453 |
| lowdim_all2349c_z | 0.368 | 0.361 | 0.386 | 0.453 |
| Feature set | Ticks | Vacc | Teach |
|---|---|---|---|
| final stack, nested | .562 | .498 | .566 |
| compact-block variant, nested | .501 | .483 | .607 |
| final judges only, nested | .484 | .464 | .589 |
| final stack (23479c pro) | .562 | .483 | .607 |
| final judges only (23479c) | .530 | .464 | .589 |
| initial stack, 3 judges (234) | .470 | .395 | .531 |