Towards Communication-Efficient Social Intelligence in Language Agents
Authors: Linxiao Gong, Yijie Xu, Tianfu Wang, Yin Wu, Yili Wang, Xingbo Yao, Huizai Yao, Xilin Xia, +2 more
Organizations: The Hong Kong University of Science and Technology (Guangzhou) · University of Science and Technology of China · The Hong Kong University of Science and Technology
Socially intelligent language agents must negotiate, coordinate, and resolve conflicting preferences while respecting the time and attention of both participants. Balancing these demands is challenging because agents must convey enough to address a partner's constraints and advance their goals without adding words that do not help the interaction. In this paper, we propose Teacher-Assisted Communication Training (TACT) to improve social goal attainment while reducing communication cost, making interactions with agents more productive and less demanding. We first characterize communication efficiency in terms of action strategy and expression, whose effects extend beyond the current utterance to the partner's response and subsequent exchanges. We design TACT to revise student-generated actions, test the revisions through partner responses, and distill useful feedback into the student. An expression specialist removes unnecessary detail while preserving the intended action, while a strategy specialist proposes alternatives that may better address the partner's constraints. To determine which revision helps, TACT samples a partner response for each candidate and selects a teacher reference by balancing local goal support against action-token cost. That reference guides on-policy distillation on the student's own generation prefixes, allowing the student to act independently at deployment. We evaluate TACT on SOTOPIA and AgentSense. On SOTOPIA, it achieves the highest Goal among the evaluated methods on All and Hard while using substantially fewer target tokens than SFT+SDPO. On AgentSense, it improves goal success over the initial student while reducing target tokens and interaction messages.
Figures & tables
Figure 1: Communication efficiency depends on the whole interaction. (A) A shorter request can require additional clarification, whereas a more informative strategy revision reaches the same goal in fewer exchanges. Dialogues are illustrative, not experimental outputs. (B) TACT lies on the empirical Goal–Cost Pareto frontier among the methods shown on both SOTOPIA-All and Hard.
Figure 2: Overview of TACT. Expression and strategy specialists propose revisions to student actions. TACT selects a reference using local interaction feedback and token cost, then conditions the teacher on this reference to guide on-policy distillation.
Method
Goal ↑
Rel. ↑
Avg ↑
Action tokens ↓
Turns ↓
SOTOPIA-All ( n=450 )
Initial student
4.327
-0.193
2.090
300.3
13.06
Concise prompt
4.691
-0.009
2.232
235.0
13.04
Teacher SFT
4.840
0.476
2.513
293.4
10.95
Vanilla OPD
4.960
0.580
2.555
275.4
11.10
Prompted OPD
5.151
0.680
2.631
275.1
10.61
Table 1: Social performance and communication cost on SOTOPIA. Avg is the seven-dimension mean. Bold and underlined values indicate the best and second-best results, respectively. Full results appear in Table 12 .
Variant
Goal ↑
Avg ↑
Action tokens ↓
Turns ↓
SOTOPIA-All ( n=450 )
Expression only
5.327
2.645
250.1
14.60
Strategy only
5.073
2.583
310.6
13.02
IG-only ranking
5.093
2.587
279.8
12.57
Token-only ranking
5.009
2.528
256.2
13.17
Random ranking
5.049
2.560
252.0
12.74
Table 2: Specialist, selection, and supervision comparisons. All/Hard panels contain 450/70 cases. TACT uses the same selected 2,970-node checkpoint as Table 1 . Full scores and configuration details appear in Table 13 ; intervals appear in Appendix F.2 .
Method
Goal (%) ↑
Rel. ↑
Action tokens ↓
Messages ↓
Initial student
48.4
0.391
899.7
14.25
Concise prompt
45.2
0.328
621.6
9.92
Vanilla OPD
54.6
0.435
841.8
13.27
TACT
54.2
0.432
789.6
12.47
Table 3: Transfer to the AgentSense two-person subset. Each method uses the same 500 instances from 100 templates. Goal is whole-goal-list success (%); Rel. is mean relationship change on [−1,1] ; Action tokens is the target agent’s generated-token count; Messages is the number of utterances from both participants. Bold and underlining mark the best and second-best displayed means, not statistical significance.
Method
Initial Qwen3.5-4B
Llama-3.1-8B
Self-play
Initial
4.327 / 3.457
4.987 / 3.400
4.576 / 3.729
Concise prompt
4.691 / 3.857
4.682 / 3.586
4.771 / 3.843
Vanilla OPD
4.960 / 3.686
5.547 / 3.371
5.251 / 2.986
TACT
5.611 / 4.371
5.807 / 3.900
6.031 / 4.214
Table 4: Goal across interaction partners. Each cell reports All / Hard Goal ( n=450/70 ). Fixed-partner columns hold the partner policy constant across methods; self-play uses the evaluated policy for both roles. Full score and cost profiles appear in Appendix F.6 .
DeepSeek-v4-pro
Kimi K2.6
GLM-5.2
Method
Goal ↑
Avg ↑
Goal ↑
Avg ↑
Goal ↑
Avg ↑
SOTOPIA-All ( n=450 )
Initial
4.327
2.090
3.080
1.211
4.244
1.944
Concise prompt
4.691
2.232
3.340
1.342
4.404
2.052
Vanilla OPD
4.960
2.555
3.636
1.768
4.769
2.370
TACT
5.611
2.711
4.191
1.874
5.251
2.500
Table 5: Cross-judge consistency on identical dialogues. Each judge scores the same 450 dialogues per method (70 Hard). Goal and seven-dimension Avg are reported; bold and underlining mark the best and second-best means. Full scores and protocols appear in Appendix F.6 .
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Student / frozen specialist backbone
Qwen3.5-4B / Qwen3.5-27B
Training partner
Student snapshot for the collection batch
Evaluation partner
Initial Qwen3.5-4B
Final dialogue judge
DeepSeek-v4-pro
OPD implementation
Sampled-token K0, upstream actor update
LoRA rank / scaling / dropout
32 / 64 / 0
Appendix
Table 6: Implementation settings. Training-generation and evaluation temperatures are distinct. Thinking generation is disabled.
Computation
Input and permitted use
Student / source partner
Own visible observations, own goal, and outward interaction; no teacher reference or partner-private goal.
Specialist proposal
Learning role’s visible messages and complete original action; its expression or strategy instruction.
Branch partner
Partner’s own visible messages after the hypothetical action; no specialist instructions or identity.
Goal-support scorer
Learning role’s before/after visible messages; its unchanged goal text as the scoring target.
OPD teacher
Original student messages plus selected reference in the final user message; original student token prefix. No branch reply.
OPD student
Original student messages and token prefix only; no selected reference.
Appendix
Table 7: Information available to each computation. “Visible history” is role-specific. The final judge has retrospective evaluation access; generated actions do not.
Metric
Range
Evaluation question
Goal
[0,10]
To what extent was the role’s stated goal achieved?
Rel.
[−5,5]
Did the interaction improve or damage relationships or standing?
Kno.
[0,10]
Did the role acquire new and relevant information?
Bel.
[0,10]
Was behavior natural and consistent with the character?
Sec.
[−10,0]
Were private information and secret intentions protected?
Rules
[−10,0]
Were moral rules or laws violated?
Appendix
Table 8: SOTOPIA dimensions and native score ranges. Higher is better for every dimension. Avg is the unnormalized arithmetic mean of all seven scores; it does not define a social-sufficiency threshold.
Role
Model and information used
Candidate specialists
Frozen Qwen3.5-27B with expression or strategy instructions.
IG scorer
Batch-frozen Qwen3.5-4B student; scores the target goal before and after the action and partner reply.
Training/branch partner
Same batch-frozen student snapshot in the partner role.
Distillation teacher
Frozen Qwen3.5-27B; conditions on the selected reference and scores original student tokens.
Evaluation partner
Fixed initial Qwen3.5-4B.
Outcome judge
DeepSeek-v4-pro; evaluates completed dialogues with the seven-dimensional SOTOPIA rubric.
Appendix
Table 9: Model roles in TACT. Specialists share a frozen backbone; the IG scorer and distillation teacher are distinct roles.
Method
Goal (%) ↑
Rel. ↑
Tokens ↓
Messages ↓
Initial
48.4 [42.4, 54.2]
0.391 [0.330, 0.452]
899.7 [844.7, 950.9]
14.25 [13.40, 15.04]
Concise prompt
45.2 [39.4, 51.0]
0.328 [0.265, 0.388]
621.6 [574.7, 667.1]
9.92 [9.18, 10.63]
Vanilla OPD
54.6 [48.8, 60.4]
0.435 [0.374, 0.492]
841.8 [791.0, 890.9]
13.27 [12.48, 14.04]
TACT
54.2 [48.0, 60.2]
0.432 [0.370, 0.487]
789.6 [736.4, 842.3]
12.47 [11.63, 13.30]
Appendix
Table 10: AgentSense means and 95% template-cluster intervals. All four panels contain 500 instances in the same 100 templates.
Comparator
Δ Goal (pp) ↑
Δ Rel. ↑
Δ Tokens ↓
Δ Messages ↓
Initial
5.8 [0.8, 10.8]
0.040 [-0.001, 0.081]
-110.1 [-156.9, -62.1]
-1.78 [-2.52, -1.02]
Concise prompt
9.0 [3.8, 14.2]
0.104 [0.063, 0.145]
168.0 [130.6, 206.2]
2.56 [1.97, 3.16]
Vanilla OPD
-0.4 [-4.8, 4.2]
-0.003 [-0.040, 0.034]
-52.2 [-91.1, -13.3]
-0.80 [-1.42, -0.18]
Appendix
Table 11: Paired AgentSense differences: TACT minus each comparator. Goal differences are percentage points; relationship differences retain the native scale. Brackets give 95% template-cluster intervals.
Method
Goal ↑
Rel. ↑
Kno. ↑
Bel. ↑
Sec. ↑
Rules ↑
Fin. ↑
Avg ↑
Tokens ↓
Turns ↓
SOTOPIA-All ( n=450 )
Initial student
4.327
-0.193
3.707
7.800
-0.442
-0.669
0.102
2.090
300.3
13.06
Concise prompt
4.691
-0.009
3.687
7.827
-0.311
-0.476
0.216
2.232
235.0
13.04
Teacher SFT
4.840
0.476
3.944
8.431
-0.213
-0.251
0.367
2.513
293.4
10.95
Vanilla OPD
4.960
0.580
4.060
8.376
-0.171
-0.224
0.307
2.555
275.4
11.10
Prompted OPD
5.151
0.680
4.122
8.413
-0.160
-0.207
0.420
2.631
275.1
10.61
Appendix
Table 12: Full social performance and communication cost on SOTOPIA. Full-dimensional results for Table 1 . Bold and underlined values indicate the best and second-best results, respectively.
Method
Goal ↑
Rel. ↑
Kno. ↑
Bel. ↑
Sec. ↑
Rules ↑
Fin. ↑
Avg ↑
Tokens ↓
Turns ↓
SOTOPIA-All ( n=450 )
Expression only
5.327
0.869
4.082
8.289
-0.147
-0.338
0.433
2.645
250.1
14.60
Strategy only
5.073
0.764
4.067
8.218
-0.162
-0.238
0.356
2.583
310.6
13.02
IG-only ranking
5.093
0.631
3.987
8.384
-0.147
-0.213
0.373
2.587
279.8
12.57
Token-only ranking
5.009
0.611
3.978
8.216
-0.162
-0.289
0.331
2.528
256.2
13.17
Random ranking
5.049
0.613
4.027
8.276
-0.136
-0.258
0.351
2.560
252.0
12.74
Appendix
Table 13: Full specialist, selection, and supervision comparisons. Full-dimensional results for Table 2 . TACT uses the same selected 2,970-node checkpoint and original evaluation cohort as Table 1 ; checkpoint details appear in Appendix D.2 . All populated rows use complete 450/70 panels after missing-case recovery. No reference retains candidate selection but removes selected-reference conditioning from the scoring teacher; its checkpoint uses 2,818 training nodes, 74 updates, and learning rate 10−5 . Intervals for the full-method and single-specialist rows appear in Appendix F.2 .
Method
Goal ↑
Rel. ↑
Kno. ↑
Avg ↑
Tokens ↓
Turns ↓
SOTOPIA-All
Initial student
[3.90, 4.77]
[-0.48, 0.09]
[3.45, 3.96]
[1.94, 2.24]
[282.7, 318.7]
[12.30, 13.82]
Concise prompt
[4.20, 5.19]
[-0.32, 0.30]
[3.43, 3.95]
[2.07, 2.39]
[222.4, 247.9]
[12.33, 13.75]
Teacher SFT
[4.36, 5.32]
[0.19, 0.75]
[3.66, 4.22]
[2.38, 2.65]
[276.2, 310.5]
[10.26, 11.66]
Vanilla OPD
[4.46, 5.46]
[0.27, 0.89]
[3.78, 4.34]
[2.41, 2.70]
[257.0, 294.3]
[10.29, 11.93]
SFT+SDPO
[4.77, 5.79]
[1.21, 1.77]
[4.05, 4.54]
[2.66, 2.93]
[381.4, 523.0]
[13.65, 15.14]
Appendix
Table 14: 95% intervals for complete-panel means. Full point estimates appear in Tables 12 and 13 .
Comparator
Δ Goal ↑
Δ Tokens ↓
Δ Turns ↓
SOTOPIA-All
Initial student
+1.284 [0.947, 1.631]
-19.7 [-33.8, -5.5]
+1.158 [0.573, 1.738]
Concise prompt
+0.920 [0.616, 1.231]
+45.6 [33.6, 57.6]
+1.182 [0.609, 1.742]
Teacher SFT
+0.771 [0.456, 1.107]
-12.8 [-26.7, 1.0]
+3.269 [2.669, 3.873]
Vanilla OPD
+0.651 [0.331, 0.976]
+5.2 [-8.0, 18.9]
+3.122 [2.540, 3.720]
SFT+SDPO
+0.336 [0.027, 0.647]
-155.1 [-234.9, -107.2]
-0.182 [-0.805, 0.413]
Appendix
Table 15: Paired differences: main-table TACT minus each comparator. Entries are mean differences [95% scenario-cluster interval]. Positive Goal and negative cost differences favor TACT. Comparators retain their recorded training pipelines.
Social performance ↑
Communication cost ↓
LR / nodes
Goal
Rel.
Kno.
Bel.
Sec.
Rules
Fin.
Avg
Tokens
Turns
SOTOPIA-All: development ( n=83/84 , respectively)
10−6 / 3,118
5.07
0.33
3.66
8.14
-0.30
-0.36
0.37
2.42
251.6
12.22
5×10−6 / 3,606
5.32
0.51
3.90
8.29
-0.20
-0.25
0.38
2.56
293.5
13.67
SOTOPIA-All: main-table reference ( n=450 )
10−5 / 2,970
5.611
0.818
4.171
8.356
-0.278
-0.229
0.527
2.711
280.6
14.22
Appendix
Table 16: Development endpoints and the selected main-table checkpoint. Development rows use baseline-paired valid cases from the 90-case panel. The 10−5 rows reproduce the complete 450/70-case evaluation in Table 12 . Nodes count admitted training actions at the evaluated checkpoint, not the final size of its training run. Panels have different evaluation coverage and checkpoint selection; no cross-panel ranking is applied.
Run
Unique configs.
Nodes
Updates
Reference, 10−5
931
3,000
59
Reference, 5×10−6
1,129
3,606
71
Reference, 10−6
1,129
3,118
71
Appendix
Table 17: Training exposure. Reference-conditioned runs at the reported training endpoints. A selected training node, a dialogue configuration, and an optimizer update are distinct units.
Stage
Expression
Strategy
Proposed
6,624 (100.0%)
6,624 (100.0%)
Scorable
6,039 (91.2%)
5,665 (85.5%)
Eligible
2,405 (36.3%)
1,370 (20.7%)
Selected
2,036 (30.7%)
934 (14.1%)
Used for OPD
2,036 (30.7%)
934 (14.1%)
Share of used references
68.6%
31.4%
Appendix
Table 18: Specialist proposal funnel and OPD adoption. Same 2,970-node checkpoint as the main results. Scorable requires both the original and candidate branches to be scorable. Funnel percentages use each specialist’s 6,624 proposals; the final row instead uses all 2,970 references actually used for updates. Counts describe supervision frequency.
Diagnostic
Count
Percentage
No eligible candidate
3,755 / 6,725
55.8%
Exactly one eligible candidate
2,165 / 6,725
32.2%
Both candidates eligible
805 / 6,725
12.0%
Disagreement with IG-only ranking
165 / 805
20.5%
Disagreement with shortest-eligible ranking
315 / 805
39.1%
Appendix
Table 19: Selection opportunities and ranking disagreements. Eligibility percentages use all 6,725 target actions. Ranking disagreements use only the 805 dual-eligible nodes and count a disagreement when the selected candidate falls outside the alternative rule’s best set; tied optima are not disagreements.
Candidate group
Pairs
Nodes
Δ Goal [95% interval]
All comparable candidates
351
218
+0.103 [ −0.012 , +0.223 ]
Positive IG gain
187
149
+0.147 [ +0.017 , +0.277 ]
Upper quartile of positive gains
45
38
+0.244 [ −0.073 , +0.605 ]
Remaining positive gains
142
122
+0.116 [ −0.021 , +0.255 ]
Appendix
Table 20: Candidate–student IG gains and complete-continuation outcomes. Exploratory groups within the random 300-node cohort. Goal changes compare a teacher candidate with the original student action; brackets give 95% scenario-cluster intervals. Counts are paired comparisons, not independent training runs; one node may contribute two pairs.
Selection
Δ Goal [95% interval]
Δ Tokens
Δ Turns
Randomly sampled nodes ( n=153 )
Shortest legal action
−0.023 [ −0.159 , +0.114 ]
−17.03
−0.157
IG-only, eligible
+0.023 [ −0.101 , +0.145 ]
−9.20
+0.020
Random, eligible
+0.025 [ −0.103 , +0.149 ]
−10.35
−0.026
TACT selection
+0.020 [ −0.106 , +0.144 ]
−9.19
+0.065
Large teacher–teacher IG gaps ( n=43 )
Appendix
Table 21: Selector outcomes on common complete training nodes. Differences are relative to freshly continuing the original student action. Goal brackets are 95% scenario-cluster intervals; cost columns are mean differences. Positive Goal and negative cost differences favor the selector.
Comparison
Pairs
Nodes
Spearman ρ [95% interval]
Randomly sampled nodes
Candidate minus student
351
218
+0.030 [ −0.088 , +0.144 ]
Eligible candidate minus student
230
193
+0.084 [ −0.048 , +0.212 ]
Expression minus strategy
163
163
−0.001 [ −0.163 , +0.163 ]
Large teacher–teacher IG gaps
Candidate minus student
111
68
+0.002 [ −0.195 , +0.202 ]
Appendix
Table 22: IG–Goal rank associations in the two continuation cohorts. Spearman correlations use paired differences, with fixed expression-minus-strategy direction for teacher–teacher comparisons. Brackets give pointwise 95% scenario-cluster intervals; exploratory comparisons are not multiplicity-adjusted. Each row uses the complete pairs required for that comparison.
Table 23: Sensitivity to the interaction partner. The same four methods are compared with the initial Qwen3.5-4B partner, Llama-3.1-8B-Instruct, and a partner using the evaluated policy (self-play). Each block uses the same 450 scenario–character configurations, including 70 Hard cases, and DeepSeek-v4-pro scoring. Fixed-initial results are reproduced from Tables 1 and 12 . Bold and underlining mark the best and second-best means among the four methods within each partner–subset block.
Method
Goal ↑
Rel. ↑
Kno. ↑
Bel. ↑
Sec. ↑
Rules ↑
Fin. ↑
Avg ↑
SOTOPIA-All ( n=450 ) — DeepSeek-v4-pro
Initial
4.327
-0.193
3.707
7.800
-0.442
-0.669
0.102
2.090
Concise prompt
4.691
-0.009
3.687
7.827
-0.311
-0.476
0.216
2.232
Vanilla OPD
4.960
0.580
4.060
8.376
-0.171
-0.224
0.307
2.555
TACT
5.611
0.818
4.171
8.356
-0.278
-0.229
0.527
2.711
SOTOPIA-All ( n=450 ) — Kimi K2.6
Appendix
Table 24: Full cross-judge score panels. Seven native-scale SOTOPIA dimensions and their arithmetic mean for the identical dialogues in Table 5 . Best and second-best displayed means are marked separately within each judge–subset block; markings do not imply statistical significance.
Multi-agent systems (MAS) built on large language models are typically organized around roles, pipelines, and turn schedules, while the content that agents pass to one another is often left as unconstrained natural language. However, this free-form communication can rapidly inflate token usage, consume the shared context window, and ultimately affect both system performance and inference cost. We analyze five common inter-agent communication strategies across two MAS topologies, finding that no fixed strategy is universally optimal. Instead, effective inter-agent messages consistently preserve action-centered information needed by downstream agents. Building on this, we propose the PACT (Protocolized Action-state Communication and Transmission), which treats inter-agent communication as a public state-update problem and projects each raw agent output into a compact action-state record before it enters shared history. Across different MAS topologies, PACT consistently improves the performance-cost trade-off, achieving comparable or stronger task performance with substantially fewer tokens. The gains extend to production coding harnesses: PACT lifts OpenHands' resolve rate at -10% tokens-per-resolved, and is resolve-neutral on SWE-agent while halving input tokens. Our code is publicly available at https://github.com/iNLP-Lab/PACT.
Large language models (LLMs) excel in structured tasks but struggle with dynamic social interactions, where success requires long-term goal coordination and rapid adaptation. Current methods often apply uniform goal-based rewards to every utterance, overlooking the specificity of objectives at each dialogue turn and failing to account for the rationale of potential strategies. Inspired by the Theory of Planned Behavior, we propose the Think-Strategy-Response (TSR) framework, which decomposes social dialogue into two hierarchical stages: high-level strategic planning and low-level linguistic execution. To optimize TSR, we introduce Linearized Hierarchical Reinforcement Learning with Variance-Gated Rewards (LHRL-VGR), a novel algorithm that dynamically routes rewards - balancing goal completion and strategy adherence - based on the variance of goal achievement scores. Experiments on the SOTOPIA benchmark show that our approach fine-tunes a Qwen2.5-7B agent to surpass the GPT-4o baseline by 7.32% in goal completion success, demonstrating state-of-the-art performance in multi-agent social negotiation tasks.
Xiaofeng Wang, Kakam Chong, Shuai Xiao +9
Independent Researcher · Alibaba Group · Shanghai Jiao Tong University
Social intelligence, the ability to navigate complex interpersonal interactions, presents a fundamental challenge for language agents. Training such agents via reinforcement learning requires solving the credit assignment problem: determining how individual utterances contribute to multi-turn dialogue outcomes. Existing approaches directly employ language models to distribute episode-level rewards, yielding attributions that are retrospective and lack theoretical grounding. We propose SAVOIR (ShApley Value fOr SocIal RL), a novel principled framework grounded in cooperative game theory. Our approach combines two complementary principles: expected utility shifts evaluation from retrospective attribution to prospective valuation, capturing an utterance's strategic potential for enabling favorable future trajectories; Shapley values ensure fair credit distribution with axiomatic guarantees of efficiency, symmetry, and marginality. Experiments on the SOTOPIA benchmark demonstrate that SAVOIR achieves new state-of-the-art performance across all evaluation settings, with our 7B model matching or exceeding proprietary models including GPT-4o and Claude-3.5-Sonnet. Notably, even large reasoning models consistently underperform, suggesting social intelligence requires qualitatively different capabilities than analytical reasoning.
Xiachong Feng, Yi Jiang, Xiaocheng Feng +9
The University of Hong Kong · Harbin Institute of Technology · Harbin Institute of Technology, Shenzhen