Whose Voice Survives the Summary? A Voice-Retention Audit of LLM Employee Listening
Organizations: Technical University of Munich · Massachusetts Institute of Technology · Leuphana University Lüneburg
Abstract
Organizations increasingly route employee feedback to leaders through large language model (LLM) summaries, an unaudited layer that silences already-spoken voice. We introduce a Voice Retention / Representation Ratio metric for representational bias in summarization and apply it to a bilingual (English/German) corpus of 2,586 free-text responses from a global professional service company. First, employees supply criticism more reliably than praise (withholding praise is 82 times more common). Second, across 45 leader-summaries the pipeline filters by popularity, not sentiment: criticism survives, yet a concern voiced once is dropped 86% of the time, with short and German-only content lost on the same axis (theme retention 0.14 vs 0.74; German directional). Controlling for frequency, sentiment has no independent effect; the harm is prevalence-driven, which sentiment-only audits miss. A targeted prompt recovers only named themes. We contribute the metric, field evidence, and a disaggregated voice-retention card.
Figures & tables
| Voice type | n | Share |
|---|---|---|
| Appreciative | 2,072 | 44.3% |
| Promotive | 1,897 | 40.6% |
| Prohibitive | 585 | 12.5% |
| Descriptive-neutral | 67 | 1.4% |
| Perfunctory | 52 | 1.1% |
| Config | Retention [95% CI] | Salience- weighted |
|---|---|---|
| exec_terse | 0.61 [0.55, 0.67] | 0.77 |
| balanced | 0.71 [0.66, 0.76] | 0.87 |
| preserve | 0.70 [0.65, 0.75] | 0.83 |
| Predictor | Coefficient | p |
|---|---|---|
| log(source units) | +1.61 | 0.001 |
| Critical theme | +0.22 | 0.55 |