Right Answers, Wrong States: Hidden Information Failures in Multi-Agent Collaboration
Organizations: Xi’an Jiaotong University · National University of Singapore · Yunnan University · Agency for Science, Technology and Research (A*STAR), Singapore
Abstract
Multi-agent systems are often judged by whether they reach the correct answer. This can miss a distinct failure: collaboration may leave behind a corrupted information state even when the immediate decision is correct. We call this an off-query failure. To study this failure in collaborative decision support, we introduce OffQuery, which separately evaluates evidence verification (T1), shared-state reconstruction (T2), and task resolution (T3) in two representative high-stakes settings: healthcare and disaster response. Across GPT, Gemini, and Qwen models, standard collaboration shows much stronger task performance than state reliability. Averaged over 21 model--setting combinations, task resolution reaches 64.7%, while evidence verification and state reconstruction reach only 14.3% and 43.1%. We trace this gap to selective information use: current queries often bypass corrupted facts, which become consequential when later tasks require them. We further introduce ReGround, which resolves conflicting evidence, verifies shared facts, reconstructs a trusted state, and reasons over that state. Across seven models from three families, ReGround improves all three capabilities in every evaluated setting, with average relative gains of 309.0%, 82.9%, and 17.6% on T1, T2, and T3. Reliable collaboration therefore requires both a correct decision and a reliable shared state for future reasoning.
Figures & tables
| Settings | Healthcare-S | Healthcare-I | Disaster Response | ||||||
|---|---|---|---|---|---|---|---|---|---|
| T1 | T2 | T3 | T1 | T2 | T3 | T1 | T2 | T3 | |
| Qwen3-8B | |||||||||
| Oracle | .621 | .882 | .675 | .642 | .887 | .586 | .613 | .867 | .710 |
| Stand. | .140 -77.5% | .317 -64.1% | .256 -62.1% | .155 -75.9% | .311 -64.9% | .256 -56.3% | .129 -79.0% | .495 -42.9% | .559 -21.3% |
| Ours | .392 +180.0% | .561 +77.0% | .462 +80.5% | .392 +152.9% | .552 +77.5% | .390 +52.3% | .656 +408.5% | .707 +42.8% | .581 +3.9% |
| Qwen3-32B | |||||||||
| Settings | Dist. | Unrel. P | Unrel. S | Healthcare-S | Healthcare-I | ||||
|---|---|---|---|---|---|---|---|---|---|
| T1 | T2 | T3 | T1 | T2 | T3 | ||||
| Oracle | ✗ | ✗ | ✗ | .653 | .882 | .943 | .680 | .887 | .910 |
| Distributed | ✓ | ✗ | ✗ | — | .580 -34.2% | .878 -6.9% | — | .583 -34.3% | .797 -12.4% |
| Unrel. Shared | ✓ | ✗ | ✓ | — | .301 -65.9% | .708 -24.9% | — | .296 -66.6% | .671 -26.3% |
| Unrel. Private | ✓ | ✓ | ✗ | .152 -76.7% | .546 -38.1% | .772 -18.1% | .140 -79.4% | .560 -36.9% | .714 -21.5% |
| Standard | ✓ | ✓ | ✓ | .152 -76.7% | .273 -69.0% | .669 -29.1% | .140 -79.4% | .287 -67.6% | .627 -31.1% |
| Baselines | GPT-4o | Gemini-3.5-Flash | ||||
|---|---|---|---|---|---|---|
| D1 | D2 | D3 | D1 | D2 | D3 | |
| Voting | .232 | .080 | .591 | .324 | .172 | .527 |
| Robin | .603 | .578 | .742 | .801 | .831 | .602 |
| Manager | .658 | .562 | .602 | .871 | .865 | .591 |
| Instructor | .444 | .395 | .548 | .867 | .845 | .602 |
| Self-Reflect | .411 | .319 | .763 | .776 | .744 | .613 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | # Inst. | Context | # Agents | |
|---|---|---|---|---|
| Words. | Sents. | |||
| Healthcare -S | 723 | 94.19 | 6.48 | 7.44 |
| Healthcare -I | 587 | 94.49 | 6.54 | 7.55 |
| Disaster Response | 93 | 148.03 | 6.67 | 5.55 |
| Judges | Oracle | GPT-4o | Gemini-3.5-Flash | Qwen3-32B | |||
|---|---|---|---|---|---|---|---|
| Stand. | Ours | Stand. | Ours | Stand. | Ours | ||
| GPT-4o | .885 | .280 | .730 | .443 | .816 | .424 | .658 |
| Gemini-3.5-Flash | .884 | .223 | .708 | .404 | .803 | .380 | .626 |
| Qwen3-235B-A22B | .880 | .219 | .725 | .406 | .809 | .380 | .634 |
| Pairs | Oracle | GPT-4o | Gemini-3.5-Flash | Qwen3-32B | Avg. | |||
|---|---|---|---|---|---|---|---|---|
| Stand. | Ours | Stand. | Ours | Stand. | Ours | |||
| Healthcare S | ||||||||
| GPT vs. Gemini | .096 | .146 | .141 | .167 | .126 | .164 | .152 | .142 |
| GPT vs. Qwen | .093 | .143 | .130 | .159 | .116 | .156 | .132 | .133 |
| Gemini vs. Qwen | .091 | .114 | .134 | .144 | .123 | .137 | .147 | .127 |
| Healthcare I | ||||||||
| Settings | Agreement | Cohen’s |
|---|---|---|
| Human 1 vs. Human 2 | 0.92 | 0.84 |
| Human 1 vs. GPT-4o | 0.88 | 0.76 |
| Human 2 vs. GPT-4o | 0.84 | 0.68 |
| Model | Provider | Model Identifier | Parameters | Access |
|---|---|---|---|---|
| Qwen3-8B ( Yang et al., 2025 ) | Alibaba | qwen3-8b | 8B | API |
| Qwen3-32B ( Yang et al., 2025 ) | Alibaba | qwen3-32b | 32B | API |
| Qwen3-235B-A22B ( Yang et al., 2025 ) | Alibaba | qwen3-235b-a22b | 235B / 22B active | API |
| GPT-4o ( Hurst et al., 2024 ) | OpenAI | gpt-4o | Not disclosed | API |
| GPT-5-Mini ( Singh et al., 2025 ) | OpenAI | gpt-5-mini | Not disclosed | API |
| GPT-5 ( Singh et al., 2025 ) | OpenAI | gpt-5 | Not disclosed | API |
| Settings | Healthcare S | Healthcare I | Disaster Response | ||||||
|---|---|---|---|---|---|---|---|---|---|
| T1 | T2 | T3 | T1 | T2 | T3 | T1 | T2 | T3 | |
| Qwen3-8B | |||||||||
| Oracle | .621 | .882 | .675 | .642 | .887 | .586 | .613 | .867 | .710 |
| Stand. | .140 -77.5% | .579 -34.4% | .390 -42.2% | .155 -75.9% | .576 -35.1% | .390 -33.4% | .129 -79.0% | .718 -17.2% | .559 -21.3% |
| Ours | .392 +180.0% | .766 +32.3% | .505 +29.5% | .392 +152.9% | .781 +35.6% | .465 +19.2% | .656 +408.5% | .814 +13.4% | .645 +15.4% |
| Qwen3-32B | |||||||||
| Settings | T1 | T2 | T3 |
|---|---|---|---|
| Vanilla | .545 | .728 | .661 |
| .413 -24.2% | .645 -11.4% | .611 -7.6% | |
| .452 -17.1% | .696 -4.4% | .654 -1.1% | |
| .558 +2.4% | .736 +1.1% | .671 +1.5% | |
| .496 -9.0% | .710 -2.5% | .638 -3.5% | |
| Oracle | 1.00 +83.5% | .813 +11.7% | .837 +26.6% |
| Settings | Healthcare -S | Healthcare -I | Disaster Response | ||||||
|---|---|---|---|---|---|---|---|---|---|
| # Input | # Output | # Total | # Input | # Output | # Total | # Input | # Output | # Total | |
| GPT-4o | |||||||||
| Standard | 46.3 | 5.7 | 52.1 | 63.6 | 7.8 | 71.5 | 26.9 | 4.4 | 31.2 |
| Chain | 19.9 | 5.2 | 25.0 | 23.3 | 6.2 | 29.5 | 13.3 | 3.5 | 16.8 |
| Tree | 20.2 | 5.2 | 25.4 | 24.0 | 6.4 | 30.3 | 13.6 | 3.6 | 17.2 |
| Voting | 3.5 | 1.5 | 5.0 | 3.6 | 1.5 | 5.1 | 3.6 | 1.4 | 5.0 |