When Upstream Messages Override Correct Answers: A Controlled Study of Multi-Agent LLM Collaboration
Organizations: University of Science and Technology of China · Qwen Business Unit of Alibaba · National University of Singapore
Abstract
Multi-agent LLM systems rely on message passing among specialized agents to accomplish complex tasks. However, an upstream agent may provide useful information or an incorrect answer that causes a downstream agent to override a correct answer supported by its own evidence. Prior work has not clearly separated the benefits of communication from the damage caused by incorrect messages. We study this problem with controlled experiments across five benchmarks and five receivers, keeping the downstream task and evidence fixed while comparing answers under three conditions: no message, the upstream agent's original message, or a message with the opposite conclusion. Our experiments reveal three key findings. First, messages often help when the downstream agent would otherwise answer incorrectly. Second, messages can also hurt: when the downstream agent would answer correctly without a message, an incorrect upstream message changes the answer in up to 32% of cases. Third, in 94% of audited harmful cases, the downstream agent copies the upstream's specific wrong answer--a pattern we term answer substitution. Removing unreliable messages recovers part of the lost accuracy, suggesting that communication should be selective based on upstream reliability and the evidence already available to the downstream agent.
Figures & tables
| Component | BIRD ( ) | LBM ( ) | L2W ( ) |
|---|---|---|---|
| A: Message removal | |||
| B: Receiver replacement | |||
| C: Total |
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| Benchmark | Task type | Upstream err. | Indep. evidence | Metric | |
|---|---|---|---|---|---|
| BIRD ( Li et al., 2023 ) | SQL generation | 150 | 61.3% | schema desc. | exec. acc. |
| L2W ( Ho et al., 2020 ) | Multi-hop QA | 120 | 13.3% | retrieved pass. | token F1 |
| LBM ( Trivedi et al., 2022 ) | Multi-hop QA | 160 | 43.8% | retrieved pass. | token F1 |
| HotpotQA ( Yang et al., 2018 ) | Multi-hop QA | 200 | 50.0% | gold paragraphs | token F1 |
| DROP ( Dua et al., 2019 ) | Reading comp. | 120 | 18.3% | passage + question | exact match |
| Receiver can solve without message? | Message value with evidence | Interpretation | |
|---|---|---|---|
| Yes (score ) | 18 | pp | Substitution |
| No (score ) | 69 | pp | Reasoning helps |
| All upstream-wrong items | 87 | pp | — |
| Behavior | BIRD | LBM | HQA | DROP |
|---|---|---|---|---|
| Copies upstream error | 84 | 66 | 38 | 86 |
| Ignores error, correct | 7 | 4 | 37 | 0 |
| Can solve, still follows | 9 | 6 | 23 | 9 |
| Cannot solve, diverges | 0 | 21 | 0 | 5 |
| Diverges, incorrect | 0 | 3 | 0 | 0 |
| Diverges, correct | 0 | 0 | 2 | 0 |
| Condition | Shown (%) | Hidden (%) | pp | |
|---|---|---|---|---|
| Without independent evidence | 86.0 | 14.2 | ||
| With independent evidence | 86.8 | 63.6 | ||
| Receiver | With evidence | No evidence | (w/o ev.) | (w/ ev.) | [95% CI] | c w | ||
|---|---|---|---|---|---|---|---|---|
| Shown | Hidden | Shown | Hidden | |||||
| gpt-4o-mini | 86.8 | 63.6 | 86.0 | 14.2 | 5 | |||
| deepseek-v3.2 | 90.8 | 70.8 | 90.0 | 14.2 | [45,66] | 3 | ||
| kimi-k2.6 | 90.8 | 83.3 | 90.8 | 21.7 | [52,71] | 7 | ||
| glm-5 | 92.5 | 90.8 | 90.8 | 20.0 | [61,78] | 2 | ||
| qwen3.6-plus | 92.5 | 94.2 | 90.0 | 24.2 | [58,76] | 6 | ||
| Benchmark | Precision (%) | Recall (%) | Items flagged |
|---|---|---|---|
| BIRD ( =150) | 91.8 | 60.9 | 61 |
| LBM ( =160) | 81.1 | 42.9 | 37 |
| L2W ( =120) | 35.0 | 43.8 | 20 |
| Intervention | BIRD ( =150, fl.=61) | LBM ( =160, fl.=37) | L2W ( =120, fl.=20) | |||
|---|---|---|---|---|---|---|
| Acc | 95% CI | Acc | 95% CI | Acc | 95% CI | |
| Default (no detection) | .400 | [.320, .480] | .539 | [.467, .609] | .742 | [.671, .807] |
| No message (same receiver) | .460 | [.380, .540] | .553 | [.482, .624] | .730 | [.656, .801] |
| Rerun (same receiver) | .447 | [.367, .527] | .553 | [.481, .623] | .755 | [.686, .819] |
| CoVe | .407 | [.327, .487] | .535 | [.462, .606] | .746 | [.675, .813] |
| LLM-judge repair | .407 | [.327, .487] | .542 | [.470, .614] | .746 | [.676, .813] |
| Component | BIRD ( ) | LBM ( ) | L2W ( ) |
|---|---|---|---|
| A: Message removal | |||
| B: Receiver replacement | |||
| C: Total |
| Receiver | Family | Solo acc. (%) | vs gpt ( pp) | Recovery ( pp) | |
| Same-family replacements (OpenAI): | |||||
| gpt-4.1-mini | OpenAI | 44.0 | |||
| gpt-4o | OpenAI | 48.7 | |||
| Cross-family replacements (capability-matched, pp): | |||||
| deepseek-v3.2 | DeepSeek | 40.7 | |||
| glm-4.5-air | Zhipu | 42.0 | |||
| Receiver | Bench | (w/o ev.) pp | (w/ ev.) pp | pp | 95% CI | |
|---|---|---|---|---|---|---|
| gpt-4o-mini | BIRD | 150 | ||||
| gpt-4o-mini | LBM | 160 | ||||
| gpt-4o-mini | L2W | 120 | ||||
| gpt-4o-mini | HQA | 200 | ||||
| deepseek-v3.2 | BIRD | 150 | ||||
| deepseek-v3.2 | LBM | 160 |
| Receiver | Benchmark | (w/ ev.) pp | 95% CI | |||
|---|---|---|---|---|---|---|
| glm-5 | BIRD | 150 | ||||
| kimi-k2.6 | BIRD | 150 | ||||
| qwen3.6-plus | BIRD | 150 | .191 | |||
| deepseek-v3.2 | BIRD | 150 | .319 | |||
| gpt-4o-mini | BIRD | 150 | .769 | |||
| kimi-k2.6 | LBM | 160 |
| Receiver | Bench | Bypass correct | Bypass wrong | ||
|---|---|---|---|---|---|
| qwen3.6-plus | BIRD | 96 | 16.7% | 54 | 16.7% |
| glm-5 | BIRD | 87 | 26.4% | 63 | 6.3% |
| kimi-k2.6 | BIRD | 75 | 32.0% | 75 | 8.0% |
| deepseek-v3.2 | BIRD | 70 | 25.7% | 80 | 15.0% |
| gpt-4o-mini | BIRD | 58 | 20.7% | 92 | 15.2% |
| qwen3.6-plus | LBM | 121 | 10.7% | 39 | 12.8% |
| Receiver | Bench | c w | pp | 95% CI | w c | ||
|---|---|---|---|---|---|---|---|
| gpt-4o-mini | BIRD | 58 | 12 | 92 | 14 | ||
| gpt-4o-mini | LBM | 77 | 9 | 83 | 23 | ||
| gpt-4o-mini | L2W | 67 | 1 | 53 | 29 | ||
| gpt-4o-mini | HQA | 159 | 21 | 41 | 8 | ||
| deepseek-v3.2 | BIRD | 70 | 18 | 80 | 12 | ||
| deepseek-v3.2 | LBM | 76 | 11 | 84 | 28 |
| Upstream | Benchmark | Ups. corr. | (w/ ev.) pp | 95% CI | pp |
|---|---|---|---|---|---|
| gpt-4o-mini | BIRD | 38.7% | |||
| gpt-4o-mini | LBM | 56.2% | |||
| gpt-5.4 | HQA | 84.5% | |||
| gpt-5.4 | LBM | 66.9% | |||
| kimi-k2.6 | HQA | 76.0% | |||
| kimi-k2.6 | LBM | 45.6% |
| Upstream wrong | Upstream correct | |||
|---|---|---|---|---|
| Benchmark | 95% CI | 95% CI | ||
| LBM | ||||
| HotpotQA | ||||
| L2W | ||||
| Without evidence | With evidence | |||||
|---|---|---|---|---|---|---|
| Benchmark | Shown | Hidden | Shown | Hidden | ||
| BIRD | 31.3 | 4.7 | 40.0 | 38.7 | ||
| LBM | 52.5 | 14.6 | 53.9 | 46.2 | ||
| L2W | 72.3 | 25.7 | 74.1 | 52.4 | ||
| HQA | 67.4 | 33.1 | 67.2 | 72.9 | ||
| Benchmark | Original (chars) | Reversed (chars) | Ratio | |
|---|---|---|---|---|
| BIRD | 150 | 1.05 | ||
| LBM | 160 | 1.03 | ||
| DROP | 120 | 0.95 | ||
| L2W | 120 | 1.04 | ||
| HotpotQA | 200 | 0.74 |
| Bench | Upstream | No msg | Orig | Rev | Rev Orig | Rev No msg | |
|---|---|---|---|---|---|---|---|
| LBM | Wrong | 70 | 11.0 | 8.0 | 50.2 | ||
| LBM | Correct | 90 | 68.2 | 90.3 | 50.9 | ||
| HQA | Wrong | 100 | 57.4 | 37.0 | 66.7 | ||
| HQA | Correct | 100 | 90.8 | 96.0 | 57.0 | ||
| L2W | Wrong | 16 | 17.1 | 18.5 | 64.5 | ||
| L2W | Correct | 104 | 54.0 | 82.7 | 10.6 |
| Benchmark | Coherent | Evidence-grounded | No artifacts | |
|---|---|---|---|---|
| BIRD | 150 | 99.3% | 100.0% | 98.0% |
| LBM | 160 | 90.0% | 98.8% | 100.0% |
| HQA | 200 | 96.3% | 100.0% | 100.0% |
| DROP | 120 | 93.3% | 100.0% | 100.0% |
| L2W | 120 | 92.5% | 100.0% | 100.0% |
| All | 750 | 94.4% | 99.7% | 99.6% |
| Benchmark | Token-F1 | ROUGE-L | Citation preservation | |
|---|---|---|---|---|
| BIRD | 150 | 0.73 | 0.68 | 50% |
| LBM | 475 | 0.62 | 0.50 | 71% |
| HotpotQA | 457 | 0.60 | 0.44 | 63% |
| Natural-error override | Injected-conflict override | |||
|---|---|---|---|---|
| Receiver | No CoT | CoT | No CoT | CoT |
| gpt-4o-mini | 100% (9/9) | 100% (9/9) | 89% (8/9) | 67% (6/9) |
| kimi-k2.6 | 93% (14/15) | 56% (10/18) | 93% (14/15) | 67% (12/18) |
| deepseek-v3.2 | — | — | 90% | 73% |
| Benchmark | Receiver | c w | McNemar | F1 | |
|---|---|---|---|---|---|
| LBM | gpt-4o-mini | 48 | 12 | pp | |
| LBM | deepseek-v3.2 | 51 | 13 | pp | |
| LBM | kimi-k2.6 | 49 | 3 | pp | |
| HQA | gpt-4o-mini | 81 | 9 | pp | |
| HQA | deepseek-v3.2 | 80 | 10 | pp | |
| HQA | kimi-k2.6 | 83 | 2 | pp |
| Benchmark | Receiver | |||||
|---|---|---|---|---|---|---|
| LBM | gpt-4o-mini | 48 | 36 | 12 | 0 | 0 |
| LBM | deepseek-v3.2 | 51 | 38 | 13 | 0 | 0 |
| LBM | kimi-k2.6 | 49 | 46 | 3 | 0 | 0 |
| HQA | gpt-4o-mini | 81 | 72 | 9 | 0 | 0 |
| HQA | deepseek-v3.2 | 80 | 70 | 10 | 0 | 0 |
| HQA | kimi-k2.6 | 83 | 81 | 2 | 0 | 0 |
| Full sample | Upstream-wrong items | ||||||
|---|---|---|---|---|---|---|---|
| Receiver | pp | 95% CI | c w / w c | pp | c w | w c | McNemar |
| gpt-4o-mini | 34 / 14 | 34 | 5 | ||||
| deepseek-v3.2 | 9 / 5 | 9 | 0 | ||||
| kimi-k2.6 | 4 / 1 | 4 | 0 | ||||
| Receiver | Condition | c w | Rate | |
|---|---|---|---|---|
| gpt-4o-mini | Matched hidden | 64 | 1 | 1.6% |
| Teammate | 64 | 17 | 26.6% | |
| Unverified tool | 64 | 16 | 25.0% | |
| Unlabeled | 64 | 18 | 28.1% | |
| Evidence-priority | 64 | 16 | 25.0% | |
| deepseek-v3.2 | Matched hidden | 64 | 1 | 1.6% |
| Dimension | Qu et al. | Cho et al. | Xie et al. | Ours |
|---|---|---|---|---|
| Setting | Multi-agent discussion | Simulated herd | Single-model context | Pipeline handoff |
| Evidence control | None | None | Parametric vs. context | Fixed gold evidence |
| Message manip. | Observe only | Majority injection | Context injection | Show/hide/reverse |
| Upstream errors | Natural | Simulated majority | Constructed | Natural |
| Causal granularity | Aggregate | Aggregate | Aggregate | Per-item paired |
| Trace analysis | No | No | No | 60 annotated CoT |
| Paper name | Role | Provider | API identifier | Temp. | Max tok. |
|---|---|---|---|---|---|
| gpt-4o-mini | Upstream/receiver | OpenAI | gpt-4o-mini | 0 | 4096 |
| gpt-5.4 | Upstream/receiver | OpenAI | gpt-5.4-0305-global | 0 | 4096 |
| deepseek-v3.2 | Receiver | DeepSeek | deepseek-v3.2 | 0 | 4096 |
| kimi-k2.6 | Upstream/receiver | Moonshot | kimi-k2.6 | 0 | 4096 |
| glm-5 | Receiver | ZhiPu | glm-5 | 0 | 4096 |
| qwen3.6-plus | Receiver | Alibaba | qwen3.6-plus | 0 | 4096 |