Organizations: Ming Hsieh Department of Electrical and Computer Engineering University of Southern California, Los Angeles, CA 90089, USA · Thomas Lord Department of Computer Science University of Southern California, Los Angeles, CA 90089, USA
Multi-agent communication aims to help agents benefit from one another's information. Yet improvements in system performance leave a fundamental ambiguity: do they reflect effective communication, a favorable agent architecture, or simply additional reasoning? Because communication methods are commonly evaluated within the systems they were designed for, these factors are difficult to disentangle. Final accuracy further merges corrected errors and corrupted answers into a single outcome, obscuring how communication changes decisions. We introduce Independent--Communicate--Revise (ICR), a controlled framework that evaluates communication as answer revision following independent reasoning. ICR fixes initial reasoning trajectories, measures correction and preservation conditional on both agents' initial correctness, and uses a no-message revision control to quantify gains beyond additional reasoning. Across four reasoning benchmarks, our audit of textual and latent communication reveals that similar aggregate accuracy can conceal substantially different revision behaviors. Compared with transmitting answers alone, full reasoning increases correction while reducing preservation on all four benchmarks, so richer messages amplify beneficial and harmful influence alike. Receiver-policy comparisons on MedQA and GPQA-D further show that a structured verification policy shifts every channel toward greater preservation and lower correction, while its effect on selectivity varies across channels and tasks. These findings challenge treating communication quality as an intrinsic property of a channel. ICR therefore recenters evaluation on selective revision, providing a unified framework for examining how message content and receiver policies jointly produce benefits and harms.
Figures & tables
Figure 1: Same final success, different intermediate contributions. ICR audits correction and preservation.
Figure 2: ICR overview. (a) Frozen initial trajectories support controlled communication and matched no-message revision. (b) Offline auditing separates four initial-correctness strata; each rate measures the fraction of correct revised answers within its stratum.
(a) Qwen3-4B
MedQA
ARC-C
GSM8K
GPQA-D
Condition
Acc. †
CR
PR
SI
Acc. †
CR
PR
SI
Acc. †
CR
PR
SI
Acc. †
CR
PR
SI
Keep Initial
–
0.00
100.00
50.00
–
0.00
100.00
50.00
–
0.00
100.00
50.00
–
0.00
100.00
50.00
Initial answers
69.33
–
–
–
93.91
–
–
–
94.62
–
–
–
54.04
–
–
–
No Message
70.28
6.90
98.28
52.59
93.93
8.16
92.86
50.51
94.74
13.04
92.39
52.72
54.58
14.04
96.49
55.26
Answer Only
–
38.79
80.17
59.48
–
38.78
63.27
51.02
–
39.13
70.65
54.89
–
56.14
63.16
59.65
Table 1: Communication audit under Critical Evaluation (%). Acc. † denotes full-set accuracy estimated under assumptions about unobserved revision outcomes; Initial answers reports measured accuracy. Estimation assumptions are detailed in Appendix G . Bold marks the highest value per metric and dataset among revision conditions within each model.
(a) Both-incorrect recovery (SR)
Condition
MedQA
ARC-C
GSM8K
GPQA-D
No Message
2.52
1.52
1.80
2.31
Full Text
0.46
0.30
0.60
1.39
StateBridge
0.46
0.30
0.90
2.55
LatentMAS
1.61
0.91
0.60
0.93
(b) Both-correct preservation (SCR)
Table 2: Recovery and preservation by initial sender–receiver correctness (%). Rates are measured on retained questions with at least one initially incorrect agent among the three. Answer Only is evaluated only on mixed-correctness pairs and is omitted here.
Figure 3: Qwen3-4B Correction–preservation profiles. Markers show point estimates; the dashed line indicates CR+PR=100% , corresponding to SI=50% . The receiver prompt is fixed, while message content and delivery vary.
Dataset
Δ CR
Δ PR
Δ SI
95% CI
MedQA
+41.38
−37.93
+1.72
[−4.31,7.76]
ARC-C
+41.84
−28.57
+6.63
[−0.51,13.27]
GSM8K
+11.96
−18.48
−3.26
[−12.50,6.52]
GPQA-D
+18.42
−15.79
+1.32
[−6.14,8.77]
Table 3: Full Text minus Answer Only, in percentage points. 95% CIs for Δ SI use 10,000 paired question-cluster bootstrap resamples.
Figure 4: Qwen3-4B Answer Only versus random delivery at matched PR. (a–d) CR for Answer Only and a reference mixing Full Text and No Message to match its PR. (e) Answer Only minus reference CR, with 95% paired question-cluster bootstrap intervals.
Critical Evaluation
Structured Verification
Policy difference
Dataset
Channel
CR
PR
SI
CR
PR
SI
Δ SI
95% CI
MedQA
No Message
6.90
98.28
52.59
3.45
99.14
51.29
−1.29
[−3.88,1.29]
Full Text
80.17
42.24
61.21
56.90
67.24
62.07
+0.86
[−5.17,6.90]
StateBridge
54.31
71.55
62.93
27.59
85.34
56.47
−6.47
[−11.64,−1.29]
LatentMAS
69.83
32.76
51.29
43.97
78.45
61.21
+9.91
[1.29,18.53]
GPQA-D
No Message
14.04
96.49
55.26
7.02
97.37
52.19
−3.07
[−7.02,0.88]
Table 4: Receiver-policy comparison on mixed-correctness directions with Qwen3-4B. CR, PR, and SI are percentages. Policy differences are computed as Structured Verification minus Critical Evaluation and reported in percentage points. The 95% intervals for Δ SI use 10,000 paired question-cluster bootstrap resamples.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Items
k=3
k=2
k=1
k=0
Initial acc.
MedQA
300
180
26
32
62
69.33
ARC-C
1165
1067
32
17
49
93.91
GSM8K
1319
1225
23
23
48
94.62
GPQA-D
198
80
24
33
61
54.04
HumanEval+
164
136
10
4
14
87.80
Appendix
Table 5: Initial three-agent correctness patterns and initial-answer accuracy (%). k denotes the number of initially correct agents. Initial accuracy is pooled over the three agents. ARC-C counts reflect the seven structural exclusions. GPQA-D phase-1 statistics are complete.
Dataset
CR
PR
SR
SCR
MedQA
0/116 (0.00%)
0/116 (0.00%)
4/436 (0.92%)
0/52 (0.00%)
ARC-C
2/98 (2.04%)
2/98 (2.04%)
4/328 (1.22%)
0/64 (0.00%)
GSM8K
4/92 (4.35%)
4/92 (4.35%)
0/334 (0.00%)
0/46 (0.00%)
GPQA-D
12/114 (10.53%)
12/114 (10.53%)
18/432 (4.17%)
0/48 (0.00%)
HumanEval+
9/28 (32.14%)
9/28 (32.14%)
20/92 (21.74%)
0/20 (0.00%)
Appendix
Table 6: Directions involving at least one initial trajectory flagged as non-terminating. This flag is not identical to an unparsed or semantically incorrect answer; GSM8K contains non-EOS trajectories with parsed answers.
CR
PR
SI
Method
All
Excl.
All
Excl.
All
Excl.
Δ SI
No Message
14.04
9.00
96.49
96.00
55.26
52.50
-2.76
Full Text
74.56
72.00
47.37
41.00
60.96
56.50
-4.46
StateBridge
64.04
61.00
51.75
47.00
57.89
54.00
-3.89
LatentMAS
61.40
56.00
50.88
48.00
56.14
52.00
-4.14
Appendix
Table 7: GPQA-D sensitivity to excluding the nine questions with at least one truncated initial trajectory. All and Excl. denote the original and filtered populations. CR and PR denominators decrease from 114 to 100 each. Metrics are percentages; Δ SI is in percentage points. The main results retain all questions.
Independent
Retained revisions
All-correct revisions
Dataset
N
No EOS
Unparsed
N
No EOS
Unparsed
N
No EOS
Unparsed
MedQA
900
1
1
2,880
1
1
102
0
0
ARC-C
3,495
2
2
2,352
9
9
832
0
0
GSM8K
3,957
2
0
2,256
2
0
938
0
0
GPQA-D
594
13
13
2,832
12
12
495
2
2
HumanEval+
492
12
12
672
48
49
96
1
1
Appendix
Table 8: Record counts, non-termination, and parsing failures. Revision counts pool No Message, Full Text, StateBridge, and LatentMAS under Critical Evaluation. Retained and observed all-three-correct questions form disjoint populations; imputed outcomes are excluded. No all-correct observations are available for the reported MedQA StateBridge condition. No EOS and Unparsed are overlapping indicators, not mutually exclusive categories.
Dataset
Condition
Accret [95% CI]
CR [95% CI]
PR [95% CI]
MedQA
No Message
25.69 [20.69, 30.97]
6.90 [2.14, 12.10]
98.28 [95.54, 100.00]
Full Text
27.22 [21.25, 33.34]
80.17 [72.06, 88.14]
42.24 [32.00, 52.90]
StateBridge
27.78 [21.67, 34.17]
54.31 [42.62, 65.45]
71.55 [61.70, 81.03]
LatentMAS
24.72 [19.72, 29.86]
69.83 [60.20, 79.09]
32.76 [24.53, 40.77]
ARC-C
No Message
27.89 [21.93, 34.01]
8.16 [2.94, 14.82]
92.86 [86.73, 97.87]
Full Text
30.27 [23.64, 37.93]
80.61 [71.15, 89.77]
34.69 [24.00, 45.92]
Appendix
Table 9: Retained-subset accuracy, CR, and PR (%), with marginal 95% percentile intervals from 2,000 question-cluster bootstrap resamples. Between-condition differences are assessed separately using paired bootstrap intervals.
Dataset
Channel
N
Both ✓
Gain
Loss
Both ×
Net
MedQA
Full Text
720
110
86
75
449
+11
StateBridge
720
144
56
41
479
+15
LatentMAS
720
98
80
87
455
-7
ARC-C
Full Text
588
102
76
62
348
+14
StateBridge
588
112
64
52
360
+12
LatentMAS
588
107
59
57
365
+2
Appendix
Table 10: Matched outcomes against No Message over retained directions. Gain denotes No Message wrong/channel correct; loss denotes the reverse. These are answer-correctness transitions, not proof of sender-answer adoption.
Dataset
Condition
N
Prior
Sender
Third
Invalid
MedQA
No Message
720
693
11
16
0
Full Text
720
496
220
3
1
StateBridge
720
588
128
4
0
LatentMAS
720
503
207
10
0
ARC-C
No Message
588
556
15
13
4
Full Text
588
429
157
1
1
Appendix
Table 11: Answer-identity categories on retained directions. Categories are assigned in the order Invalid, Prior, Sender, and Third. Sender therefore denotes a match that differs from the receiver’s initial answer. Third also includes changed valid outputs when the sender’s initial answer is unparsed. Program comparisons use string identity rather than functional equivalence.
Sender matches
Dataset
Condition
CR
PR
Valid change rate
MedQA
Answer Only
45/116
23/116
68/232 (29.31%)
Full Text
93/116
67/116
160/232 (68.97%)
ARC-C
Answer Only
38/98
36/98
74/196 (37.76%)
Full Text
79/98
64/98
143/196 (72.96%)
GSM8K
Answer Only
36/92
23/92
64/184 (34.78%)
Appendix
Table 12: Answer Only and Full Text on shared mixed-correctness directions under Critical Evaluation. Sender matches report counts over all events in each stratum, using Prior-before-Sender precedence. Valid change rates pool CR and PR and exclude unparsed revised outputs. Each GPQA-D condition has one unparsed revised output; the other conditions shown have none.
Dataset
Condition
Prompt tok.
Output tok.
Message size
Revision (s)
MedQA
No Message
912.2
884.5
–
19.93
Full Text
1438.2
978.7
518.1 tok.
23.20
StateBridge
983.2
958.3
320.0 KiB
21.73
LatentMAS
919.2
582.6
Not recovered
14.45
ARC-C
No Message
709.7
807.5
–
20.73
Full Text
1195.8
760.1
478.2 tok.
16.40
Appendix
Table 13: Logged per-record means, not a controlled throughput comparison. Prompt-token accounting does not include all latent-prefix computation. Construction costs are not uniformly available. LatentMAS payload zeros in the source export are suppressed because they conflict with the documented KV-prefix transmission.
Condition
Accret
CR
PR
SI
SR
SCR
No Message
42.26
16/28
26/28
75.00
9/92
20/20
Full Text
45.24
21/28
26/28
83.93
9/92
20/20
StateBridge
43.45
19/28
27/28
82.14
8/92
19/20
LatentMAS
39.88
24/28
17/28
73.21
6/92
20/20
Appendix
Table 14: HumanEval+ supplementary results. Accuracy and SI are percentages; the remaining columns show success counts. All conditions evaluate the same 168 retained directions.
Dataset
Condition
Retained
Skipped sample
Imputed
Coverage
Estimate
MedQA
No Message
720
34/34
1046
41.89
70.28
Full Text
720
34/34
1046
41.89
70.89
StateBridge
720
-
1080
40.00
71.11
LatentMAS
720
34/34
1046
41.89
69.89
ARC-C
No Message
588
208/208
6194
11.39
93.93
Full Text
588
208/208
6194
11.39
94.13
Appendix
Table 15: Provisional full-set accuracy estimates, not measured full-set accuracy. Skipped-sample entries show correct/observed outcomes on all-correct triples. Coverage and estimates are percentages. The sampled directions need not constitute a probability sample of the unobserved directions. Exported interval bounds are not reproduced as confidence intervals; see text.
Dataset
Channel
Direction
N
CR
PR
SR
SCR
MedQA
No Message
A → B
120
1/18
21/22
2/72
8/8
A → C
120
1/18
19/20
2/74
8/8
B → A
120
0/22
18/18
3/72
8/8
B → C
120
1/20
18/18
3/72
10/10
C → A
120
2/20
18/18
1/74
8/8
C → B
120
3/18
20/20
0/72
10/10
Appendix
Table 16: Per-direction outcomes. Conditional columns show successes/denominator.
Dataset
α
Ref. CR (%)
Δ CR
95% CI
MedQA
0.3231
30.57
+8.22
[−2.99,19.86]
ARC-C
0.5088
45.02
−6.25
[−19.78,7.86]
GSM8K
0.5405
33.61
+5.52
[−9.85,19.01]
GPQA-D
0.6786
55.11
+1.03
[−14.52,15.57]
Appendix
Table 17: Answer Only minus the random-delivery reference at matched PR. Reference CR is a percentage; differences and intervals are in percentage points. 95% CIs use 10,000 paired question-cluster bootstrap resamples.
Dataset
D(pCE)
D(pSV)
I
95% CI for I
MedQA
+1.72
−5.60
−7.33
[−15.09,0.00]
GPQA-D
−3.07
+2.19
+5.26
[−4.39,14.91]
Appendix
Table 18: StateBridge minus Full Text in SI under each receiver policy, and the change in this contrast. All values are in percentage points. Contrasts and interactions are computed before rounding. 95% CIs use 10,000 paired question-cluster bootstrap resamples.
Dataset
nCR/nPR
Δ CR
95% CI
Δ PR
95% CI
MedQA
116/116
+41.38
[30.17,52.59]
−37.93
[−49.14,−26.72]
ARC-C
98/98
+41.84
[28.57,54.08]
−28.57
[−40.82,−16.33]
GSM8K
92/92
+11.96
[0.00,23.91]
−18.48
[−31.52,−5.43]
GPQA-D
114/114
+18.42
[8.77,28.07]
−15.79
[−26.32,−6.14]
Appendix
Table 19: Full Text minus Answer Only on shared mixed-correctness directions. Differences and intervals are in percentage points. The 95% CIs use 10,000 paired question-cluster bootstrap resamples, retaining associated directions and compared conditions together.
Dataset
Channel
G(c,CE)
G(c,SV)
ΔG
95% CI for ΔG
MedQA
Full Text
+8.62
+10.78
+2.16
[−3.88,8.19]
StateBridge
+10.34
+5.17
−5.17
[−11.21,0.86]
LatentMAS
−1.29
+9.91
+11.21
[2.16,20.26]
GPQA-D
Full Text
+5.70
+6.14
+0.44
[−7.89,8.77]
StateBridge
+2.63
+8.33
+5.70
[−2.19,13.60]
LatentMAS
+0.88
+7.02
+6.14
[−2.63,15.35]
Appendix
Table 20: Communication increments relative to the policy-specific No Message control and their changes from CE to SV. All values are in percentage points. The 95% CIs use 10,000 paired question-cluster bootstrap resamples, jointly resampling all four condition–policy cells.
Contrast
Condition
Estimate
95% CI
ΔSI
No Message
−1.35
[−4.91,2.23]
Full Text
−3.15
[−9.95,3.60]
StateBridge
+2.25
[−4.68,9.09]
ΔG
Full Text
−1.80
[−9.91,6.25]
StateBridge
+3.60
[−3.90,11.24]
Interaction I
StateBridge versus Full Text
+5.41
[−4.51,15.39]
Appendix
Table 21: GPQA-D receiver-policy contrasts after jointly excluding events with a truncated revision in any of the six evaluated condition–policy combinations. The common subset contains 111 CR and 111 PR events. All values are in percentage points. Intervals use paired question-cluster bootstrap resampling, retaining associated directions and compared conditions together.
Dataset
Questions
k=0
k=1
k=2
k=3
Initial acc. (%)
MedQA
300
40
24
22
214
78.89
GPQA-D
198
57
25
29
87
57.91
Appendix
Table 22: Qwen3-8B initial correctness. k is the number of initially correct agents. Initial accuracy is measured over all three agents; questions with k=1 or k=2 contribute mixed-correctness directions.
Dataset
Condition
CR [95% CI]
PR [95% CI]
SI
MedQA
No Message
7.61 [1.35, 15.12]
93.48 [86.90, 98.72]
50.54
Answer Only
31.52 [20.54, 42.86]
85.87 [76.32, 94.23]
58.70
Full Text
77.17 [66.04, 86.96]
46.74 [33.70, 60.47]
61.96
StateBridge
60.87 [48.89, 72.73]
76.09 [66.67, 84.91]
68.48
LatentMAS
73.91 [63.54, 83.72]
41.30 [30.30, 52.56]
57.61
GPQA-D
No Message
14.81 [7.55, 22.92]
92.59 [86.76, 97.46]
53.70
Appendix
Table 23: Qwen3-8B audit results under Critical Evaluation (%). CR and PR denominators are 92 each on MedQA and 108 each on GPQA-D. Marginal 95% intervals use 10,000 question-cluster bootstrap resamples. Bold marks the highest SI point estimate per dataset.
Dataset
Condition
Δ SI
95% CI
MedQA
Answer Only
+8.15
[2.17,14.39]
Full Text
+11.41
[1.97,20.75]
StateBridge
+17.93
[9.90,25.71]
LatentMAS
+7.07
[−1.92,16.07]
GPQA-D
Answer Only
+4.63
[−2.31,11.57]
Full Text
+3.70
[−4.41,12.07]
Appendix
Table 24: Qwen3-8B SI differences relative to No Message, in percentage points. Intervals use 10,000 paired question-cluster bootstrap resamples, preserving associated directions and compared conditions together. Differences are computed before rounding.
Multi-agent LLM systems rely on message passing among specialized agents to accomplish complex tasks. However, an upstream agent may provide useful information or an incorrect answer that causes a downstream agent to override a correct answer supported by its own evidence. Prior work has not clearly separated the benefits of communication from the damage caused by incorrect messages. We study this problem with controlled experiments across five benchmarks and five receivers, keeping the downstream task and evidence fixed while comparing answers under three conditions: no message, the upstream agent's original message, or a message with the opposite conclusion. Our experiments reveal three key findings. First, messages often help when the downstream agent would otherwise answer incorrectly. Second, messages can also hurt: when the downstream agent would answer correctly without a message, an incorrect upstream message changes the answer in up to 32% of cases. Third, in 94% of audited harmful cases, the downstream agent copies the upstream's specific wrong answer--a pattern we term answer substitution. Removing unreliable messages recovers part of the lost accuracy, suggesting that communication should be selective based on upstream reliability and the evidence already available to the downstream agent.
Yaxin Gong, Gangyi Zhang, Chongming Gao +7
University of Science and Technology of China · Qwen Business Unit of Alibaba · National University of Singapore
Multi-agent AI systems can improve answer selection by allowing different language models to exchange reasoning traces, revise initial predictions, and support a final decision. However, such communication may also introduce reliability risks: reasoning from one agent can correct another agent's mistake, but it can also mislead an agent that was initially correct. This paper studies reliable multi-agent AI communication through reasoning exchange and runtime answer revision. We develop a framework in which agents first answer multiple-choice questions independently, then share reasoning traces and revise their decisions. We conduct numerical experiments where we evaluate whether this process improves accuracy, produces more positive than negative answer transitions, and remains effective across domains such as cybersecurity, networking, and general knowledge. The results help identify when multi-agent reasoning improves reliability and when it may propagate errors.
Shahnewaz Karim Sakib, Anindya Bijoy Das
University of Tennessee at Chattanooga, TN 37403, USA · The University of Akron, OH 44325, USA
Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-relevant information. End-task performance alone also cannot reveal whether an observed effect depends on message presence, content generated for the evaluated example, or information supplied by a separate agent. We introduce a causal audit that applies controlled message replacements at the boundary where the sender-produced representation enters the receiver. Four message settings support five measurements of encoded sender information, receiver sensitivity to message presence and identity, the task value of example-specific content, and the additional value supplied by a separate agent. We apply the audit to latent relay with Qwen3-4B and Qwen3-8B on GSM8K, ARC-C, and MATH-500. On GSM8K, the Qwen3-4B overall performance effect of -1.00 percentage point decomposes into a -6.17-point effect retained by an other-example message and a +5.17-point effect attributable to example-specific content; both component directions reverse at 8B. On MATH-500, the Qwen3-4B gain of 15.00 points comprises 8.33 points retained by an other-example message and 6.67 points attributable to example-specific content, while the 8B gain is dominated by the former component. Self-substitution comparisons further show that example-specific content and other-agent value are distinct. These results show that aggregate accuracy does not identify how a latent message affects the receiver and motivate controlled message comparisons as a standard evaluation for latent communication.
Huixiang Zhang, Mahzabeen Emu
Georgia Institute of Technology · North Ave · Atlanta, GA 30332 USA +3