Towards Communication-Efficient Social Intelligence in Language Agents
Organizations: The Hong Kong University of Science and Technology (Guangzhou) · University of Science and Technology of China · The Hong Kong University of Science and Technology
Abstract
Socially intelligent language agents must negotiate, coordinate, and resolve conflicting preferences while respecting the time and attention of both participants. Balancing these demands is challenging because agents must convey enough to address a partner's constraints and advance their goals without adding words that do not help the interaction. In this paper, we propose Teacher-Assisted Communication Training (TACT) to improve social goal attainment while reducing communication cost, making interactions with agents more productive and less demanding. We first characterize communication efficiency in terms of action strategy and expression, whose effects extend beyond the current utterance to the partner's response and subsequent exchanges. We design TACT to revise student-generated actions, test the revisions through partner responses, and distill useful feedback into the student. An expression specialist removes unnecessary detail while preserving the intended action, while a strategy specialist proposes alternatives that may better address the partner's constraints. To determine which revision helps, TACT samples a partner response for each candidate and selects a teacher reference by balancing local goal support against action-token cost. That reference guides on-policy distillation on the student's own generation prefixes, allowing the student to act independently at deployment. We evaluate TACT on SOTOPIA and AgentSense. On SOTOPIA, it achieves the highest Goal among the evaluated methods on All and Hard while using substantially fewer target tokens than SFT+SDPO. On AgentSense, it improves goal success over the initial student while reducing target tokens and interaction messages.
Figures & tables
| Method | Goal | Rel. | Avg | Action tokens | Turns |
|---|---|---|---|---|---|
| SOTOPIA-All ( ) | |||||
| Initial student | 4.327 | -0.193 | 2.090 | 300.3 | 13.06 |
| Concise prompt | 4.691 | -0.009 | 2.232 | 235.0 | 13.04 |
| Teacher SFT | 4.840 | 0.476 | 2.513 | 293.4 | 10.95 |
| Vanilla OPD | 4.960 | 0.580 | 2.555 | 275.4 | 11.10 |
| Prompted OPD | 5.151 | 0.680 | 2.631 | 275.1 | 10.61 |
| Variant | Goal | Avg | Action tokens | Turns |
|---|---|---|---|---|
| SOTOPIA-All ( ) | ||||
| Expression only | 5.327 | 2.645 | 250.1 | 14.60 |
| Strategy only | 5.073 | 2.583 | 310.6 | 13.02 |
| IG-only ranking | 5.093 | 2.587 | 279.8 | 12.57 |
| Token-only ranking | 5.009 | 2.528 | 256.2 | 13.17 |
| Random ranking | 5.049 | 2.560 | 252.0 | 12.74 |
| Method | Goal (%) | Rel. | Action tokens | Messages |
|---|---|---|---|---|
| Initial student | 48.4 | 0.391 | 899.7 | 14.25 |
| Concise prompt | 45.2 | 0.328 | 621.6 | 9.92 |
| Vanilla OPD | 54.6 | 0.435 | 841.8 | 13.27 |
| TACT | 54.2 | 0.432 | 789.6 | 12.47 |
| Method | Initial Qwen3.5-4B | Llama-3.1-8B | Self-play |
|---|---|---|---|
| Initial | 4.327 / 3.457 | 4.987 / 3.400 | 4.576 / 3.729 |
| Concise prompt | 4.691 / 3.857 | 4.682 / 3.586 | 4.771 / 3.843 |
| Vanilla OPD | 4.960 / 3.686 | 5.547 / 3.371 | 5.251 / 2.986 |
| TACT | 5.611 / 4.371 | 5.807 / 3.900 | 6.031 / 4.214 |
| DeepSeek-v4-pro | Kimi K2.6 | GLM-5.2 | ||||
| Method | Goal | Avg | Goal | Avg | Goal | Avg |
| SOTOPIA-All ( ) | ||||||
| Initial | 4.327 | 2.090 | 3.080 | 1.211 | 4.244 | 1.944 |
| Concise prompt | 4.691 | 2.232 | 3.340 | 1.342 | 4.404 | 2.052 |
| Vanilla OPD | 4.960 | 2.555 | 3.636 | 1.768 | 4.769 | 2.370 |
| TACT | 5.611 | 2.711 | 4.191 | 1.874 | 5.251 | 2.500 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Value |
|---|---|
| Student / frozen specialist backbone | Qwen3.5-4B / Qwen3.5-27B |
| Training partner | Student snapshot for the collection batch |
| Evaluation partner | Initial Qwen3.5-4B |
| Final dialogue judge | DeepSeek-v4-pro |
| OPD implementation | Sampled-token K0, upstream actor update |
| LoRA rank / scaling / dropout | 32 / 64 / 0 |
| Computation | Input and permitted use |
|---|---|
| Student / source partner | Own visible observations, own goal, and outward interaction; no teacher reference or partner-private goal. |
| Specialist proposal | Learning role’s visible messages and complete original action; its expression or strategy instruction. |
| Branch partner | Partner’s own visible messages after the hypothetical action; no specialist instructions or identity. |
| Goal-support scorer | Learning role’s before/after visible messages; its unchanged goal text as the scoring target. |
| OPD teacher | Original student messages plus selected reference in the final user message; original student token prefix. No branch reply. |
| OPD student | Original student messages and token prefix only; no selected reference. |
| Metric | Range | Evaluation question |
|---|---|---|
| Goal | To what extent was the role’s stated goal achieved? | |
| Rel. | Did the interaction improve or damage relationships or standing? | |
| Kno. | Did the role acquire new and relevant information? | |
| Bel. | Was behavior natural and consistent with the character? | |
| Sec. | Were private information and secret intentions protected? | |
| Rules | Were moral rules or laws violated? |
| Role | Model and information used |
|---|---|
| Candidate specialists | Frozen Qwen3.5-27B with expression or strategy instructions. |
| IG scorer | Batch-frozen Qwen3.5-4B student; scores the target goal before and after the action and partner reply. |
| Training/branch partner | Same batch-frozen student snapshot in the partner role. |
| Distillation teacher | Frozen Qwen3.5-27B; conditions on the selected reference and scores original student tokens. |
| Evaluation partner | Fixed initial Qwen3.5-4B. |
| Outcome judge | DeepSeek-v4-pro; evaluates completed dialogues with the seven-dimensional SOTOPIA rubric. |
| Method | Goal (%) | Rel. | Tokens | Messages |
|---|---|---|---|---|
| Initial | 48.4 [42.4, 54.2] | 0.391 [0.330, 0.452] | 899.7 [844.7, 950.9] | 14.25 [13.40, 15.04] |
| Concise prompt | 45.2 [39.4, 51.0] | 0.328 [0.265, 0.388] | 621.6 [574.7, 667.1] | 9.92 [9.18, 10.63] |
| Vanilla OPD | 54.6 [48.8, 60.4] | 0.435 [0.374, 0.492] | 841.8 [791.0, 890.9] | 13.27 [12.48, 14.04] |
| TACT | 54.2 [48.0, 60.2] | 0.432 [0.370, 0.487] | 789.6 [736.4, 842.3] | 12.47 [11.63, 13.30] |
| Comparator | Goal (pp) | Rel. | Tokens | Messages |
|---|---|---|---|---|
| Initial | 5.8 [0.8, 10.8] | 0.040 [-0.001, 0.081] | -110.1 [-156.9, -62.1] | -1.78 [-2.52, -1.02] |
| Concise prompt | 9.0 [3.8, 14.2] | 0.104 [0.063, 0.145] | 168.0 [130.6, 206.2] | 2.56 [1.97, 3.16] |
| Vanilla OPD | -0.4 [-4.8, 4.2] | -0.003 [-0.040, 0.034] | -52.2 [-91.1, -13.3] | -0.80 [-1.42, -0.18] |
| Method | Goal | Rel. | Kno. | Bel. | Sec. | Rules | Fin. | Avg | Tokens | Turns |
|---|---|---|---|---|---|---|---|---|---|---|
| SOTOPIA-All ( ) | ||||||||||
| Initial student | 4.327 | -0.193 | 3.707 | 7.800 | -0.442 | -0.669 | 0.102 | 2.090 | 300.3 | 13.06 |
| Concise prompt | 4.691 | -0.009 | 3.687 | 7.827 | -0.311 | -0.476 | 0.216 | 2.232 | 235.0 | 13.04 |
| Teacher SFT | 4.840 | 0.476 | 3.944 | 8.431 | -0.213 | -0.251 | 0.367 | 2.513 | 293.4 | 10.95 |
| Vanilla OPD | 4.960 | 0.580 | 4.060 | 8.376 | -0.171 | -0.224 | 0.307 | 2.555 | 275.4 | 11.10 |
| Prompted OPD | 5.151 | 0.680 | 4.122 | 8.413 | -0.160 | -0.207 | 0.420 | 2.631 | 275.1 | 10.61 |
| Method | Goal | Rel. | Kno. | Bel. | Sec. | Rules | Fin. | Avg | Tokens | Turns |
|---|---|---|---|---|---|---|---|---|---|---|
| SOTOPIA-All ( ) | ||||||||||
| Expression only | 5.327 | 0.869 | 4.082 | 8.289 | -0.147 | -0.338 | 0.433 | 2.645 | 250.1 | 14.60 |
| Strategy only | 5.073 | 0.764 | 4.067 | 8.218 | -0.162 | -0.238 | 0.356 | 2.583 | 310.6 | 13.02 |
| IG-only ranking | 5.093 | 0.631 | 3.987 | 8.384 | -0.147 | -0.213 | 0.373 | 2.587 | 279.8 | 12.57 |
| Token-only ranking | 5.009 | 0.611 | 3.978 | 8.216 | -0.162 | -0.289 | 0.331 | 2.528 | 256.2 | 13.17 |
| Random ranking | 5.049 | 0.613 | 4.027 | 8.276 | -0.136 | -0.258 | 0.351 | 2.560 | 252.0 | 12.74 |
| Method | Goal | Rel. | Kno. | Avg | Tokens | Turns |
|---|---|---|---|---|---|---|
| SOTOPIA-All | ||||||
| Initial student | [3.90, 4.77] | [-0.48, 0.09] | [3.45, 3.96] | [1.94, 2.24] | [282.7, 318.7] | [12.30, 13.82] |
| Concise prompt | [4.20, 5.19] | [-0.32, 0.30] | [3.43, 3.95] | [2.07, 2.39] | [222.4, 247.9] | [12.33, 13.75] |
| Teacher SFT | [4.36, 5.32] | [0.19, 0.75] | [3.66, 4.22] | [2.38, 2.65] | [276.2, 310.5] | [10.26, 11.66] |
| Vanilla OPD | [4.46, 5.46] | [0.27, 0.89] | [3.78, 4.34] | [2.41, 2.70] | [257.0, 294.3] | [10.29, 11.93] |
| SFT+SDPO | [4.77, 5.79] | [1.21, 1.77] | [4.05, 4.54] | [2.66, 2.93] | [381.4, 523.0] | [13.65, 15.14] |
| Comparator | Goal | Tokens | Turns |
|---|---|---|---|
| SOTOPIA-All | |||
| Initial student | +1.284 [0.947, 1.631] | -19.7 [-33.8, -5.5] | +1.158 [0.573, 1.738] |
| Concise prompt | +0.920 [0.616, 1.231] | +45.6 [33.6, 57.6] | +1.182 [0.609, 1.742] |
| Teacher SFT | +0.771 [0.456, 1.107] | -12.8 [-26.7, 1.0] | +3.269 [2.669, 3.873] |
| Vanilla OPD | +0.651 [0.331, 0.976] | +5.2 [-8.0, 18.9] | +3.122 [2.540, 3.720] |
| SFT+SDPO | +0.336 [0.027, 0.647] | -155.1 [-234.9, -107.2] | -0.182 [-0.805, 0.413] |
| Social performance | Communication cost | |||||||||
| LR / nodes | Goal | Rel. | Kno. | Bel. | Sec. | Rules | Fin. | Avg | Tokens | Turns |
| SOTOPIA-All: development ( , respectively) | ||||||||||
| / 3,118 | 5.07 | 0.33 | 3.66 | 8.14 | -0.30 | -0.36 | 0.37 | 2.42 | 251.6 | 12.22 |
| / 3,606 | 5.32 | 0.51 | 3.90 | 8.29 | -0.20 | -0.25 | 0.38 | 2.56 | 293.5 | 13.67 |
| SOTOPIA-All: main-table reference ( ) | ||||||||||
| / 2,970 | 5.611 | 0.818 | 4.171 | 8.356 | -0.278 | -0.229 | 0.527 | 2.711 | 280.6 | 14.22 |
| Run | Unique configs. | Nodes | Updates |
|---|---|---|---|
| Reference, | 931 | 3,000 | 59 |
| Reference, | 1,129 | 3,606 | 71 |
| Reference, | 1,129 | 3,118 | 71 |
| Stage | Expression | Strategy |
|---|---|---|
| Proposed | 6,624 (100.0%) | 6,624 (100.0%) |
| Scorable | 6,039 (91.2%) | 5,665 (85.5%) |
| Eligible | 2,405 (36.3%) | 1,370 (20.7%) |
| Selected | 2,036 (30.7%) | 934 (14.1%) |
| Used for OPD | 2,036 (30.7%) | 934 (14.1%) |
| Share of used references | 68.6% | 31.4% |
| Diagnostic | Count | Percentage |
|---|---|---|
| No eligible candidate | 3,755 / 6,725 | 55.8% |
| Exactly one eligible candidate | 2,165 / 6,725 | 32.2% |
| Both candidates eligible | 805 / 6,725 | 12.0% |
| Disagreement with IG-only ranking | 165 / 805 | 20.5% |
| Disagreement with shortest-eligible ranking | 315 / 805 | 39.1% |
| Candidate group | Pairs | Nodes | Goal [95% interval] |
|---|---|---|---|
| All comparable candidates | 351 | 218 | [ , ] |
| Positive IG gain | 187 | 149 | [ , ] |
| Upper quartile of positive gains | 45 | 38 | [ , ] |
| Remaining positive gains | 142 | 122 | [ , ] |
| Selection | Goal [95% interval] | Tokens | Turns |
|---|---|---|---|
| Randomly sampled nodes ( ) | |||
| Shortest legal action | [ , ] | ||
| IG-only, eligible | [ , ] | ||
| Random, eligible | [ , ] | ||
| TACT selection | [ , ] | ||
| Large teacher–teacher IG gaps ( ) | |||
| Comparison | Pairs | Nodes | Spearman [95% interval] |
| Randomly sampled nodes | |||
| Candidate minus student | 351 | 218 | [ , ] |
| Eligible candidate minus student | 230 | 193 | [ , ] |
| Expression minus strategy | 163 | 163 | [ , ] |
| Large teacher–teacher IG gaps | |||
| Candidate minus student | 111 | 68 | [ , ] |
| Method | Goal | Rel. | Kno. | Bel. | Sec. | Rules | Fin. | Avg | Tokens | Turns |
|---|---|---|---|---|---|---|---|---|---|---|
| SOTOPIA-All ( ) — Fixed initial Qwen3.5-4B | ||||||||||
| Initial | 4.327 | -0.193 | 3.707 | 7.800 | -0.442 | -0.669 | 0.102 | 2.090 | 300.3 | 13.06 |
| Concise prompt | 4.691 | -0.009 | 3.687 | 7.827 | -0.311 | -0.476 | 0.216 | 2.232 | 235.0 | 13.04 |
| Vanilla OPD | 4.960 | 0.580 | 4.060 | 8.376 | -0.171 | -0.224 | 0.307 | 2.555 | 275.4 | 11.10 |
| TACT | 5.611 | 0.818 | 4.171 | 8.356 | -0.278 | -0.229 | 0.527 | 2.711 | 280.6 | 14.22 |
| SOTOPIA-All ( ) — Fixed Llama-3.1-8B-Instruct | ||||||||||
| Method | Goal | Rel. | Kno. | Bel. | Sec. | Rules | Fin. | Avg |
|---|---|---|---|---|---|---|---|---|
| SOTOPIA-All ( ) — DeepSeek-v4-pro | ||||||||
| Initial | 4.327 | -0.193 | 3.707 | 7.800 | -0.442 | -0.669 | 0.102 | 2.090 |
| Concise prompt | 4.691 | -0.009 | 3.687 | 7.827 | -0.311 | -0.476 | 0.216 | 2.232 |
| Vanilla OPD | 4.960 | 0.580 | 4.060 | 8.376 | -0.171 | -0.224 | 0.307 | 2.555 |
| TACT | 5.611 | 0.818 | 4.171 | 8.356 | -0.278 | -0.229 | 0.527 | 2.711 |
| SOTOPIA-All ( ) — Kimi K2.6 | ||||||||