Multi-agent LLM systems rely on message passing among specialized agents to accomplish complex tasks. However, an upstream agent may provide useful information or an incorrect answer that causes a downstream agent to override a correct answer supported by its own evidence. Prior work has not clearly separated the benefits of communication from the damage caused by incorrect messages. We study this problem with controlled experiments across five benchmarks and five receivers, keeping the downstream task and evidence fixed while comparing answers under three conditions: no message, the upstream agent's original message, or a message with the opposite conclusion. Our experiments reveal three key findings. First, messages often help when the downstream agent would otherwise answer incorrectly. Second, messages can also hurt: when the downstream agent would answer correctly without a message, an incorrect upstream message changes the answer in up to 32% of cases. Third, in 94% of audited harmful cases, the downstream agent copies the upstream's specific wrong answer--a pattern we term answer substitution. Removing unreliable messages recovers part of the lost accuracy, suggesting that communication should be selective based on upstream reliability and the evidence already available to the downstream agent.
Figures & tables
Figure 1: An erroneous peer message overrides an otherwise correct, evidence-supported answer.
Figure 2: Evidence–message interaction, conditional message value, and upstream model variation.
Figure 3: Override rates across 25 cells and per-benchmark transition breakdown.
Figure 4: Cross-model conclusion reversal. (a) Corrupting a correct upstream conclusion universally lowers accuracy (11/11 cells). (b) Correcting a wrong conclusion universally restores it (11/11 cells). Gray dots = original message; colored dots = after reversal.
Figure 5: Source-label control, delayed-receipt control, and CoT override rates.
Figure 6: Error-location interventions and gated recovery actions.
Component
BIRD ( n=92 )
LBM ( n=70 )
L2W ( n=16 )
A: Message removal
+10.9
+6.7
+1.7
B: Receiver replacement
+10.9
+24.1
+22.3
C: Total
+21.7
+30.8
+24.0
Table 1: Three-way decomposition on upstream-wrong items (paired bootstrap, B=10,000 ). On upstream-correct items, message removal is harmful (Appendix Table 9 ).
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Task type
n
Upstream err.
Indep. evidence
Metric
BIRD ( Li et al., 2023 )
SQL generation
150
61.3%
schema desc.
exec. acc.
L2W ( Ho et al., 2020 )
Multi-hop QA
120
13.3%
retrieved pass.
token F1
LBM ( Trivedi et al., 2022 )
Multi-hop QA
160
43.8%
retrieved pass.
token F1
HotpotQA ( Yang et al., 2018 )
Multi-hop QA
200
50.0%
gold paragraphs
token F1
DROP ( Dua et al., 2019 )
Reading comp.
120
18.3%
passage + question
exact match
Appendix
Table 2: Benchmark overview. “Upstream err.” is the fraction of items where the upstream agent (gpt-4o-mini) produces an incorrect answer.
Receiver can solve without message?
n
Message value with evidence
Interpretation
Yes (score >0.5 )
18
−22.8 pp
Substitution
No (score ≤0.5 )
69
+19.1 pp
Reasoning helps
All upstream-wrong items
87
+10.5 pp
—
Appendix
Table 3: LBM items where kimi-k2.6 gives an incorrect upstream answer. Positive message value with independent evidence comes from items the receiver cannot solve alone; on items it can solve, the wrong conclusion causes harm.
Behavior
BIRD
LBM
HQA
DROP
Copies upstream error
84
66
38
86
Ignores error, correct
7
4
37
0
Can solve, still follows
9
6
23
9
Cannot solve, diverges
0
21
0
5
Diverges, incorrect
0
3
0
0
Diverges, correct
0
0
2
0
Appendix
Table 4: Behavioral classification for the gpt-4o-mini receiver (%). Percentages are rounded independently; “Follows upstream error” combines unconditional copying with cases where the receiver can solve alone but still follows. Column sample sizes: BIRD 92, LBM 70, HQA 100, DROP 22.
Condition
Shown (%)
Hidden (%)
τ pp
Without independent evidence
86.0
14.2
+71.8
With independent evidence
86.8
63.6
+23.2
Γ
+48.6
p<0.0001
Appendix
Table 5: DROP four-cell decomposition (gpt-4o-mini receiver, n =120). τ = message value (shown − hidden accuracy). Γ=τ(without evidence)−τ(with evidence) . DROP shows the largest Γ among all five benchmarks. Unlike HotpotQA and LBM, τ(with evidence) remains positive ( +23.2 pp): the message is net-helpful even with evidence. Independent evidence sharply reduces the marginal value of the message.
Receiver
With evidence
No evidence
τ (w/o ev.)
τ (w/ ev.)
Γ [95% CI]
c → w
Shown
Hidden
Shown
Hidden
gpt-4o-mini
86.8
63.6
86.0
14.2
+71.8
+23.2
+48.6
5
deepseek-v3.2
90.8
70.8
90.0
14.2
+75.8
+20.0
+55.8 [45,66]
3
kimi-k2.6
90.8
83.3
90.8
21.7
+69.2
+7.5
+61.7 [52,71]
7
glm-5
92.5
90.8
90.8
20.0
+70.8
+1.7
+69.2 [61,78]
2
qwen3.6-plus
92.5
94.2
90.0
24.2
+65.8
−1.7
+67.5 [58,76]
6
Appendix
Table 6: DROP cross-receiver four-cell decomposition ( n =120 per receiver). Accuracy is exact-match percentage. Γ is significantly positive for all five receivers ( p<0.0001 ); 95% CIs from paired bootstrap ( B =10,000).
Benchmark
Precision (%)
Recall (%)
Items flagged
BIRD ( n =150)
91.8
60.9
61
LBM ( n =160)
81.1
42.9
37
L2W ( n =120)
35.0
43.8
20
Appendix
Table 7: Error-detector operating characteristics (gemini-2.5-pro). Precision = fraction of flagged items that are truly upstream-wrong; recall = fraction of upstream-wrong items that are flagged. L2W has low precision because most upstream answers are correct (86.7%), so even modest false-positive rates dominate.
Intervention
BIRD ( n =150, fl.=61)
LBM ( n =160, fl.=37)
L2W ( n =120, fl.=20)
Acc
95% CI
Acc
95% CI
Acc
95% CI
Default (no detection)
.400
[.320, .480]
.539
[.467, .609]
.742
[.671, .807]
No message (same receiver)
.460
[.380, .540]
.553
[.482, .624]
.730
[.656, .801]
Rerun (same receiver)
.447
[.367, .527]
.553
[.481, .623]
.755
[.686, .819]
CoVe
.407
[.327, .487]
.535
[.462, .606]
.746
[.675, .813]
LLM-judge repair
.407
[.327, .487]
.542
[.470, .614]
.746
[.676, .813]
Appendix
Table 8: Recovery experiment: gated accuracy under different interventions with gemini-2.5-pro as the shared detector. “No message (different family)” uses a receiver from another model family without the upstream message. “Answer, then review message” lets the receiver answer independently before reviewing the peer message.
Component
BIRD ( n=150 )
LBM ( n=160 )
L2W ( n=120 )
A: Message removal
−1.3
−7.7
−21.8
B: Receiver replacement
+9.7
+17.4
+14.8
C: Total
+8.3
+9.7
−7.0
Appendix
Table 9: Three-way decomposition on all with-evidence items (unconditional on upstream correctness; paired bootstrap, B=10,000 ). Component A is negative when most upstream answers are correct, because removing the message discards correct signals. On the upstream-wrong target population, Component A is non-negative on all three benchmarks (Table 1 ). C=A+B by construction.
Table 10: Recovery by receiver identity on BIRD ( n =92 upstream-wrong items, paired bootstrap). Solo accuracy is computed on 150 BIRD items without any peer message in an independent run. Recovery = paired accuracy gain over the default gpt-4o-mini receiver with the erroneous message; the default receiver’s same-model removal baseline is +10.9 pp(Table 1 ). Same-family OpenAI replacements achieve recovery comparable to cross-family receivers at similar capability levels ( p=0.50 for same-family vs. cross-family mean difference), indicating that recovery does not require family diversity—any different model suffices.
Receiver
Bench
n
τ (w/o ev.) pp
τ (w/ ev.) pp
Γ pp
95% CI
gpt-4o-mini
BIRD
150
+26.7
+1.3
+25.3
[+17.3,+33.3]
gpt-4o-mini
LBM
160
+37.9
+7.7
+30.2
[+22.1,+38.6]
gpt-4o-mini
L2W
120
+46.6
+21.8
+24.8
[+15.2,+34.4]
gpt-4o-mini
HQA
200
+34.3
−5.7
+40.0
[+33.7,+46.6]
deepseek-v3.2
BIRD
150
+27.3
−4.0
+31.3
[+22.7,+40.0]
deepseek-v3.2
LBM
160
+38.3
+10.5
+27.8
[+20.7,+34.8]
Appendix
Table 11: Full 20-cell Γ table. τ (w/o ev.) = message value without independent evidence; τ (w/ ev.) = message value with evidence; Γ=τ(w/oev.)−τ(w/ev.) . All CIs from paired bootstrap ( B=10,000 , seed 42). All 20 cells show Γ>0 ( p<0.0001 ): independent evidence consistently reduces the marginal value of the upstream message.
Receiver
Benchmark
n
τ (w/ ev.) pp
95% CI
p
glm-5
BIRD
150
−12.7
[−19.3,−6.7]
<.001
∗∗∗
kimi-k2.6
BIRD
150
−12.0
[−18.7,−5.3]
<.001
∗∗∗
qwen3.6-plus
BIRD
150
−4.7
[−11.3,+2.0]
.191
deepseek-v3.2
BIRD
150
−4.0
[−11.3,+3.3]
.319
gpt-4o-mini
BIRD
150
+1.3
[−5.3,+8.0]
.769
kimi-k2.6
LBM
160
−10.2
[−16.0,−4.4]
<.001
∗∗∗
Appendix
Table 12: Paired bootstrap 95% CIs for τ (with evidence) ( B=10,000 resamples, seed 42). Of 20 cells, 13 are significantly negative at α=0.05 ; 6 survive Bonferroni correction ( α=0.0025 for 20 tests). ∗p<.05 ; ∗∗p<.01 ; ∗∗∗p<.001 .
Receiver
Bench
Bypass correct
P(c→w)
Bypass wrong
P(w→c)
qwen3.6-plus
BIRD
96
16.7%
54
16.7%
glm-5
BIRD
87
26.4%
63
6.3%
kimi-k2.6
BIRD
75
32.0%
75
8.0%
deepseek-v3.2
BIRD
70
25.7%
80
15.0%
gpt-4o-mini
BIRD
58
20.7%
92
15.2%
qwen3.6-plus
LBM
121
10.7%
39
12.8%
Appendix
Table 13: Per-item transitions under τ (with evidence). P(c→w) : fraction of items correct without the message that become wrong with it. P(w→c) : the reverse. On HQA, all five receivers show non-trivial correct-to-wrong rates (5.7–13.2%).
Receiver
Bench
ns
c → w
τs pp
95% CI
nf
w → c
gpt-4o-mini
BIRD
58
12
−20.7
[−31.0,−10.3]
92
14
gpt-4o-mini
LBM
77
9
−11.7
[−19.5,−5.2]
83
23
gpt-4o-mini
L2W
67
1
−1.5
[−4.5,0.0]
53
29
gpt-4o-mini
HQA
159
21
−13.2
[−18.9,−8.2]
41
8
deepseek-v3.2
BIRD
70
18
−25.7
[−35.7,−15.7]
80
12
deepseek-v3.2
LBM
76
11
−14.5
[−22.4,−6.6]
84
28
Appendix
Table 14: Conditional decomposition of τ (with evidence). ns = items solvable without message ( k=3 majority vote); τs = message value on solvable items (always ≤0 ). CIs from bootstrap ( B=10,000 ). All c → w transitions across the primary 20 cells: 274 out of 2,184 solvable items (12.5% overall). Including DROP, the total is 297/2,667 (11.1%). In 18 of 20 primary cells, the c → w rate exceeds 5%.
Upstream
Benchmark
Ups. corr.
τ (w/ ev.) pp
95% CI
Γ pp
gpt-4o-mini
BIRD
38.7%
+1.3
[−5.3,+8.0]
+25.3
gpt-4o-mini
LBM
56.2%
+7.7
[+1.3,+14.2]
+30.2
gpt-5.4
HQA
84.5%
+2.3
[−1.1,+5.9]
+43.1
gpt-5.4
LBM
66.9%
+29.3
[+21.8,+36.9]
+21.4
kimi-k2.6
HQA
76.0%
+1.0
[−2.5,+4.5]
+42.6
kimi-k2.6
LBM
45.6%
+20.4
[+13.5,+27.3]
+22.5
Appendix
Table 15: Upstream model diversity: message value with evidence and interaction. τ (w/ ev.) = message value with evidence; Γ=τ (w/o ev.) −τ (w/ ev.). All Γ values significant ( p<0.0001 ). τ (w/ ev.) is non-negative in all cells.
Figure 7: Conclusion-reversal experiment on items with an incorrect upstream answer (three QA benchmarks). Points show the accuracy difference between messages with corrected and original peer conclusions, with 95% bootstrap CIs. All six contrasts are significant (BIRD and DROP omitted; see text). A compact version appears in the main text as Figure 5 (a).
Upstream wrong
Upstream correct
Benchmark
Δ
95% CI
Δ
95% CI
LBM
+42.2
[31.6,53.1]
−39.4
[−49.3,−29.6]
HotpotQA
+22.7
[15.1,30.4]
−35.9
[−44.8,−27.6]
L2W
+46.0
[25.0,67.9]
−72.1
[−81.4,−61.9]
Appendix
Table 16: Effects of reversing the peer conclusion, with 95% bootstrap CIs (20K draws). All six contrasts on the three QA benchmarks are significant (BIRD and DROP rows omitted due to SQL-execution scoring and source-wrong label divergence; raw values retained in comments). “Upstream wrong”: corrected − original; “Upstream correct”: corrupted − original.
Without evidence
With evidence
Benchmark
Shown
Hidden
τ
Shown
Hidden
τ
BIRD
31.3
4.7
+26.7
40.0
38.7
+1.3
LBM
52.5
14.6
+37.9
53.9
46.2
+7.7
L2W
72.3
25.7
+46.6
74.1
52.4
+21.8
HQA
67.4
33.1
+34.3
67.2
72.9
−5.7
Appendix
Table 17: Accuracy (%) and message value τ for four benchmarks (gpt-4o-mini receiver). Within each evidence condition, τ is the paired per-item difference between message-shown and message-hidden accuracy. Classification uses k=3 majority vote (Appendix A.20 ). DROP is reported separately in Table 5 .
Benchmark
n
Original (chars)
Reversed (chars)
Ratio
BIRD
150
1167±400
1221±447
1.05
LBM
160
649±175
667±185
1.03
DROP
120
487±130
461±129
0.95
L2W
120
844±248
874±252
1.04
HotpotQA
200
1241±315
921±410
0.74
Appendix
Table 18: Character length of original messages and messages with reversed conclusions (mean ± std).
Bench
Upstream
n
No msg
Orig
Rev
Rev − Orig
Rev − No msg
LBM
Wrong
70
11.0
8.0
50.2
+42.2
+39.2
LBM
Correct
90
68.2
90.3
50.9
−39.4
−17.3
HQA
Wrong
100
57.4
37.0
66.7
+29.7
+9.3
HQA
Correct
100
90.8
96.0
57.0
−39.0
−33.8
L2W
Wrong
16
17.1
18.5
64.5
+46.0
+47.4
L2W
Correct
104
54.0
82.7
10.6
−72.1
−43.4
Appendix
Table 19: Accuracy (%) under three message conditions. “Rev” = reversed conclusion (correcting wrong or corrupting correct). All three conditions are from the same run; “Rev − Orig” is computed within this run. For bootstrapped CIs on paired effect sizes from the main experiment, see Table 16 . DROP and L2W have small source-wrong samples ( n =22 and n =16). L2W source-wrong values ( n =16) are mean token-F1 × 100, not binary accuracy.
Benchmark
n
Coherent
Evidence-grounded
No artifacts
BIRD
150
99.3%
100.0%
98.0%
LBM
160
90.0%
98.8%
100.0%
HQA
200
96.3%
100.0%
100.0%
DROP
120
93.3%
100.0%
100.0%
L2W
120
92.5%
100.0%
100.0%
All
750
94.4%
99.7%
99.6%
Appendix
Table 20: LLM-judge semantic audit of all 750 messages with reversed conclusions. Overall, 93.8% pass all three criteria. The 5.6% flagged as incoherent are conservative: an incoherent message is harder for the receiver to adopt, working against our finding that reversing conclusions shifts accuracy.
Benchmark
n
Token-F1
ROUGE-L
Citation preservation
BIRD
150
0.73
0.68
50%
LBM
475
0.62
0.50
71%
HotpotQA
457
0.60
0.44
63%
Appendix
Table 21: Token-level overlap between original and reversed messages (filtered to messages ≥ 20 tokens). n counts message pairs across all receiver models that generated flips on each benchmark, hence exceeding per-benchmark item counts. Citation preservation measures the fraction of named entities and numbers shared between versions.
Figure 29
Natural-error override
Injected-conflict override
Receiver
No CoT
CoT
No CoT
CoT
gpt-4o-mini
100% (9/9)
100% (9/9)
89% (8/9)
67% (6/9)
kimi-k2.6
93% (14/15)
56% (10/18)
93% (14/15)
67% (12/18)
deepseek-v3.2
—
—
90%
73%
Appendix
Table 23: CoT override rates on LBM. Parentheses show override count / eligible items. For kimi’s natural-error condition, the paired rate on the 11 items qualifying under both CoT conditions is 91% → 45%. Deepseek has no natural-error override cases on LBM in this CoT experiment (its 11 c → w items on LBM in the main design, Table 14 , do not fall in the natural-error CoT-eligible subset).
Benchmark
Receiver
n
c → w
McNemar p
Δ F1
LBM
gpt-4o-mini
48
12
<0.001
−22.3 pp
LBM
deepseek-v3.2
51
13
<0.001
−21.7 pp
LBM
kimi-k2.6
49
3
0.250
−3.0 pp
HQA
gpt-4o-mini
81
9
0.004
−8.9 pp
HQA
deepseek-v3.2
80
10
0.002
−11.5 pp
HQA
kimi-k2.6
83
2
0.500
−1.5 pp
Appendix
Table 24: Minimal-edit conclusion reversal across three receivers and three benchmarks. n = number of items answered correctly in an independent evidence-only response before either message condition, with valid minimal-edit messages ( ≥0.90 overlap). c → w = items that flip from correct (with the original message) to wrong (with the minimal edit). Δ F1 = mean F1 difference (minimal-flip − original message).
Benchmark
Receiver
n
a
b
c
d
LBM
gpt-4o-mini
48
36
12
0
0
LBM
deepseek-v3.2
51
38
13
0
0
LBM
kimi-k2.6
49
46
3
0
0
HQA
gpt-4o-mini
81
72
9
0
0
HQA
deepseek-v3.2
80
70
10
0
0
HQA
kimi-k2.6
83
81
2
0
0
Appendix
Table 25: Full 2×2 paired contingency tables for the minimal-edit control. a = correct under both messages; b = correct with original, wrong with minimal edit; c = wrong with original, correct with minimal edit; d = wrong under both. Items are selected by independent evidence-only correctness before either message condition. Because such items nearly always remain correct under the original (source-correct) message, c=0 and d=0 in every cell are structural consequences of this selection, not design constraints. Binarization threshold: token-F1 ≥0.5 .
Full sample
Upstream-wrong items
Receiver
τ pp
95% CI
c → w / w → c
τ pp
c → w
w → c
McNemar p
gpt-4o-mini
−4.0
[−6.3,−1.7]
34 / 14
−10.4
34
5
<0.001
deepseek-v3.2
−0.0
[−3.4,+3.3]
9 / 5
−5.5
9
0
0.004
kimi-k2.6
−1.2
[−3.4,+0.7]
4 / 1
−3.1
4
0
0.125
Appendix
Table 26: Matched-review control on HotpotQA ( n=400 for gpt-4o-mini, n=200 for deepseek and kimi). Both message-shown and message-hidden branches use the same two-step review prompt; the hidden branch receives a neutral placeholder instead of the peer message. τ = paired per-item accuracy difference (shown − hidden). On upstream-wrong items, c → w = items correct under hidden but wrong under shown; w → c = reverse. The gpt-4o-mini result is significant on the full sample (McNemar p=0.006 ) and highly significant on upstream-wrong items ( p<0.001 ). Deepseek is also significant ( p=0.004 ). Kimi shows the same directional pattern but lacks power.
Receiver
Condition
n
c → w
Rate
gpt-4o-mini
Matched hidden
64
1
1.6%
Teammate
64
17
26.6%
Unverified tool
64
16
25.0%
Unlabeled
64
18
28.1%
Evidence-priority
64
16
25.0%
deepseek-v3.2
Matched hidden
64
1
1.6%
Appendix
Table 27: Source-label attribution control on HotpotQA (items where the receiver is independently correct and the upstream answer is wrong). Every labeled condition produces more correct-to-wrong transitions than the hidden baseline; no pairwise comparison between labeled conditions is significant. Pooling across receivers, all labeled-vs-hidden comparisons are significant (Fisher exact p<0.001 ). At the individual-receiver level, all comparisons are significant for gpt-4o-mini and deepseek ( p<0.05 ); kimi’s low base rate (3–6 events out of 70) leaves most comparisons underpowered (teammate p=0.12 , unlabeled p=0.014 ).
Dimension
Qu et al.
Cho et al.
Xie et al.
Ours
Setting
Multi-agent discussion
Simulated herd
Single-model context
Pipeline handoff
Evidence control
None
None
Parametric vs. context
Fixed gold evidence
Message manip.
Observe only
Majority injection
Context injection
Show/hide/reverse
Upstream errors
Natural
Simulated majority
Constructed
Natural
Causal granularity
Aggregate
Aggregate
Aggregate
Per-item paired
Trace analysis
No
No
No
60 annotated CoT
Appendix
Table 28: Experimental-design comparison with closest prior work. Qu et al. study conformity in multi-agent discussion without controlling receiver evidence; Cho et al. inject simulated majorities without per-item pairing; Xie et al. study parametric-vs-contextual conflicts, not conflicts between two external inputs. Our design fixes downstream evidence and manipulates the upstream message item by item, enabling causal attribution.
Paper name
Role
Provider
API identifier
Temp.
Max tok.
gpt-4o-mini
Upstream/receiver
OpenAI
gpt-4o-mini
0
4096
gpt-5.4
Upstream/receiver
OpenAI
gpt-5.4-0305-global
0
4096
deepseek-v3.2
Receiver
DeepSeek
deepseek-v3.2
0
4096
kimi-k2.6
Upstream/receiver
Moonshot
kimi-k2.6
0
4096
glm-5
Receiver
ZhiPu
glm-5
0
4096
qwen3.6-plus
Receiver
Alibaba
qwen3.6-plus
0
4096
Appendix
Table 29: Model manifest. The conclusion editor generates messages with reversed peer conclusions; the detector identifies candidate errors in recovery experiments. All models were accessed in September 2026.
Large Language Models (LLMs) have enabled collaborative Multi-Agent (MA) systems, where interacting agents improve performance through diverse reasoning and iterative refinement. However, these systems remain vulnerable to error propagation, where early-stage information degrades downstream reasoning. To address this, we conduct a systematic analysis of inter-agent communication to identify which information drives MA performance. We find that the absence of reasoning and verification in inter-agent communication significantly degrades performance. Based on these insights, we propose Category-Aware Recovery Augmentation (technique), which enforces the presence of critical information during communication. recovers up to 86.2% of failed cases. Our results highlight the key role of information quality in effective MA collaboration. Our code is available at https://anonymous.4open.science/r/cara_mas
Yong Jin Chun, Iftekhar Ahmed
Department of Informatics University of California, Irvine Irvine, CA, 92618
Large language models are increasingly used in multi-agent systems, where they see and respond to other agents' answers. A key risk is conformity: a model may abandon its own answer simply because others agree on a different one. Prior studies show that LLMs often revise toward a majority answer, but it remains unclear whether these revisions help correct mistakes as often as they introduce new errors. In this paper, we conduct a controlled study in which an LLM first answers a question, then sees simulated peer responses before making a final decision. We manipulate two social cues: consensus structure and authority labels assigned to peers, and measure how they influence beneficial and harmful revisions. Across four open-weight LLMs and seven QA datasets, we find that peer agreement makes it much easier to mislead initially correct models than to correct initially wrong ones. Authority labels make models more likely to choose the endorsed answer, regardless of whether it is correct. More concerningly, generic reasoning interventions such as chain-of-thought and reflection do not reliably reduce harmful revision while preserving beneficial revision. These findings suggest that multi-agent LLM systems should verify peer answers rather than simply aggregate them.
Jiaming Qu, Lucheng Fu, Yibo Hu
Amazon · Georgia Institute of Technology · Illinois Institute of Technology
Multi-agent communication aims to help agents benefit from one another's information. Yet improvements in system performance leave a fundamental ambiguity: do they reflect effective communication, a favorable agent architecture, or simply additional reasoning? Because communication methods are commonly evaluated within the systems they were designed for, these factors are difficult to disentangle. Final accuracy further merges corrected errors and corrupted answers into a single outcome, obscuring how communication changes decisions. We introduce Independent--Communicate--Revise (ICR), a controlled framework that evaluates communication as answer revision following independent reasoning. ICR fixes initial reasoning trajectories, measures correction and preservation conditional on both agents' initial correctness, and uses a no-message revision control to quantify gains beyond additional reasoning. Across four reasoning benchmarks, our audit of textual and latent communication reveals that similar aggregate accuracy can conceal substantially different revision behaviors. Compared with transmitting answers alone, full reasoning increases correction while reducing preservation on all four benchmarks, so richer messages amplify beneficial and harmful influence alike. Receiver-policy comparisons on MedQA and GPQA-D further show that a structured verification policy shifts every channel toward greater preservation and lower correction, while its effect on selectivity varies across channels and tasks. These findings challenge treating communication quality as an intrinsic property of a channel. ICR therefore recenters evaluation on selective revision, providing a unified framework for examining how message content and receiver policies jointly produce benefits and harms.
Shixuan Li, Wei Yang, Peiyu Zhang +3
Ming Hsieh Department of Electrical and Computer Engineering University of Southern California, Los Angeles, CA 90089, USA · Thomas Lord Department of Computer Science University of Southern California, Los Angeles, CA 90089, USA