Large language model (LLM) agentic systems increasingly rely on models communicating with one another, yet existing uncertainty and multi-agent methods rarely estimate how a particular receiver will interpret a message before it is sent. This matters in heterogeneous systems, where capable receivers can reconstruct different tasks from the same message. We model this as a sender-receiver problem with a latent receiver type and define prospective interpretation risk (PIR): the probability that a receiver reconstructs a task other than intended. Rather than model an LLM's full input-output behaviour, we use black-box probes relating messages, intended tasks, and receiver-specific reconstructions, yielding scalable supervision while separating interpretation from downstream capability failure. Offline, heterogeneous frozen receivers provide supervision for receiver-conditioned risk and the effects of predefined mutable message features. At deployment, history induces a posterior over receiver types, guiding message revision and selection. We introduce value of interpretation information (VoII), querying for receiver information only when its expected communication benefit exceeds its cost. Our theory characterises when receiver information has decision value and bounds such queries. Empirically, interpretation-failure rates vary by 4-13x across receivers. Receiver information reduces PIR calibration error by 68% relative to a receiver-agnostic predictor, largely by correcting receiver-specific risk levels. PIR-guided revision reduces interpretation failure by 44% relative to the original message and 40% relative to a generic rewrite, mostly through a repair that helps every receiver. VoII outperforms information-gain and random querying at matched cost on the interpretation objective it optimises, lowering interpretation failure from 3.84% to 3.79% while querying 18.2% of episodes.
Figures & tables
Figure 1: Framework overview. Top: receivers read one message as different tasks (1), PIR separates interpretation from capability risk (2), and the sender infers the receiver and queries only if worthwhile (3). Bottom: offline probes train the PIR model, which picks the verified rewrite to send.
Figure 2: Interpretation failure is separate from execution failure and depends on the receiver (test split). (a) YA after a correct read ( ∘ ) and a misread ( ∙ ) of the probe. (b) YI per receiver. Bands span the lowest to the highest receiver.
Per receiver ↑
Across receivers ↑
Calibration ↓
Method
AUROC
AUPRC
AUROC
AUPRC
Brier
NLL
ECE
A0 Global prior
.500
.045
.500
.045
.0428
.180
.029
A0 Receiver prior
.500
.045
.697
.086
.0415
.168
.002
A1 Message judge
.552
.057
.540
.055
.0428
.180
.031
A2 Equiv. judge
.500
.045
.500
.045
.0428
.180
.029
A4 YA predictor
.511
.048
.530
.050
.4229
1.329
.540
Table 1: Prospective PIR prediction (test split, six-dataset means, seed 0). Ours is best on all seven metrics (bold, second best underlined, A0 excluded).
YI↓
ΔYI (ours − row)
YA↓
Condition
Gen.
Sel.
(%)
points [95% CI]
(%)
Original message
–
–
4.16
−1.85 [ −2.09 , −1.62 ] ‡
59.33
Generic rewrite
–
–
3.85
−1.54 [ −1.78 , −1.30 ] ‡
58.15
Selection only
–
✓
3.47
−1.15 [ −1.41 , −0.93 ] ‡
58.20
Posterior only
πt
✓
3.43
−1.11 [ −1.33 , −0.90 ] ‡
58.59
Guidance only
✓
–
2.29
+0.03 [ +0.01 , +0.05 ]
58.33
Table 2: Message revision on 1,579 held-out questions: ours beats every pre-registered reference on YI ( ‡ Holm p<0.001 ). Gen.: learned risk guides generation ( πt : posterior only). Sel.: it selects the sent message.
Primary cohort
Replication
YI
Gap
Net
Ident.
YI
Gap
Policy (query rate)
(%)
(%)
utility
(%)
(%)
(%)
No query (0%)
3.84
0
.9616
24.6
3.93
0
Exact 20% quota
Random (3 seeds)
3.84
3
.9614
25.7
3.93
4
IG
3.83
5
.9615
28.5
3.94
<0
Table 3: Query policies (interpretation objective, c=0.001 ). Gap: share of the no-query to true-identity YI gap that is closed. Ident.: posterior mode is correct. IG: information gain. † Secondary comparator. Bold: best deployable value.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
Defined in
General
E , P (or P ), Var
Expectation, probability, and variance
–
1{⋅}
Indicator: 1 if the condition holds and 0 otherwise
–
H(⋅) , I(⋅;⋅∣⋅)
Shannon entropy and conditional mutual information, in nats
–
DKL(⋅∥⋅)
Kullback–Leibler divergence
–
Ber(p)
Bernoulli distribution with mean p
–
Appendix
Table 4: Notation, grouped by the part of the framework that introduces each symbol. The last column gives where the symbol is defined, and – marks standard notation.
Receiver
First
Second
Last
All
First or second [95% CI]
Qwen3.6
1.28
1.61
2.33
1.74
1.45 [1.26, 1.65]
Gemma
1.59
2.49
2.57
2.22
2.04 [1.81, 2.28]
gpt-oss
3.27
2.08
1.63
2.33
2.68 [2.42, 2.95]
Nemotron
3.14
4.51
24.49
10.72
3.83 [3.57, 4.09]
Mistral
5.26
5.74
4.45
5.16
5.50 [5.15, 5.85]
Highest/lowest
4.1
3.6
15.0
6.2
3.8 [3.4, 4.3]
Appendix
Table 5: Probe error rate (%) by the position of the intended task (test split, parsed probes). Last column: probes with the intended task first or second, with question-cluster bootstrap 95% CIs.
ID
Method
Definition
A0
Base-rate priors
Global P(YI=1) and receiver-specific P(YI=1∣R) estimated on training data.
A1
Message-only ambiguity
Fixed strong judge sees m only and scores whether materially different task interpretations are possible.
A2
Semantic equivalence
Judge sees (m,z⋆,xS) and predicts whether m faithfully specifies the intended task, with no receiver information.
A3
Receiver-agnostic PIR
Same learned PIR architecture but remove eR / receiver posterior information.
A4
Action-failure predictor
Same architecture/input budget trained against YA rather than YI , evaluated against both targets.
A5
Listener/ToM
Predict P(ZR=z∣m,x,eR) over probe task identities and use 1−P(ZR=z⋆) as risk.
Appendix
Table 6: Baseline definitions.
Figure 3: Interpretation failure depends on the receiver and the dataset (test split). (a) YI (%) per receiver and dataset on a log colour scale. Boxes mark the least risky receiver in each dataset, and the bottom row gives the ratio of the highest to the lowest receiver. (b) Misread pairs split by the answer check. Labels give the share of misreads that pass.
Marginal rates (%)
Off-diagonal cells (%)
Dataset
Pairs
YI [95% CI]
YA [95% CI]
misread, pass
read, fail
FreebaseQA
6,550
2.5 [1.7, 3.4]
24.0 [21.7, 26.4]
1.8
23.3
WebQSP
1,095
5.5 [2.5, 8.6]
54.1 [47.5, 60.7]
1.8
50.3
GrailQA
6,430
7.9 [6.5, 9.4]
89.6 [87.9, 91.3]
0.8
82.4
SQ-WD
7,090
5.2 [4.1, 6.4]
61.9 [59.3, 64.4]
2.2
58.8
PopQA
7,255
3.3 [2.4, 4.2]
75.7 [73.5, 77.9]
0.8
73.2
Appendix
Table 7: Interpretation ( YI ) and execution ( YA ) failures per dataset (test split, pooled over five receivers, 95% question-level normal intervals).
Per receiver ↑
Across receivers ↑
Calibration ↓
Method
AUROC
AUPRC
AUROC
AUPRC
Brier
NLL
ECE
A3 (no receiver)
.729
.128
.701
.115
.0414
.168
.032
A3 + recv. intercept
.729
.128
.792
.158
.0401
.156
.015
A3 + recv. calib.
.728
.128
.796
.158
.0399
.155
.012
A3 + recv. map ∗
.728
.128
.797
.160
.0398
.154
.010
Ours
.732
.135
.800
.164
.0398
.154
.010
Appendix
Table 8: Receiver levels or interaction (Experiment 2, test split, six-dataset means, seed 0). Maps are refitted on the validation questions used by the full model ( ∗ on all validation data). Difference: ours minus A3 + recv. calib., with paired question-group bootstrap 95% CIs.
Figure 4: Receiver information calibrates the predicted risk (test split, seed 0 in (a, b), three-seed means in (c)). (a) Within-receiver ECE per dataset. (b) Mean predicted versus measured YI in each of the 30 receiver–dataset cells. (c) AUROC across versus within receivers, with arrows from A3 to ours.
Dataset
Method
AUROC [95% CI]
AUPRC
Brier ↓
NLL ↓
ECE ↓
FreebaseQA
A0 global base rate
0.500
0.025
0.0246
0.118
0.024
A0 per-receiver base rate
0.500
0.025
0.0238
0.105
0.003
A1 message-only judge
0.529 [0.481, 0.578]
0.029
0.0246
0.118
0.024
A2 equivalence judge
0.500
0.025
0.0246
0.118
0.024
A3 receiver-agnostic PIR
0.690 [0.645, 0.730]
0.085
0.0243
0.113
0.024
A4 action-failure predictor
0.558 [0.518, 0.599]
0.032
0.115
0.377
0.234
Appendix
Table 9: Experiment 2 results per dataset (test split, seed 0, with metrics computed within each receiver and macro-averaged). AUROC has a paired 95% bootstrap CI. Bold: best value per dataset among learned methods.
Variant
AUROC (pooled) ↑
AUROC (macro) ↑
Brier ↓
NLL ↓
ECE (macro) ↓
FreebaseQA
Full model (with eR )
0.810
0.698
0.0234
0.100
0.005
Receiver removed (A3)
0.663
0.691
0.0243
0.113
0.024
Receiver shuffled
0.661
0.691
0.0243
0.113
0.024
Δ full − A3
+ 0.147
+ 0.008
− 0.0009
− 0.013
95% CI
[0.124, 0.170]
[ − 0.0012, − 0.0007]
[ − 0.015, − 0.011]
Appendix
Table 10: Receiver-information ablation (test split, three-seed means). Δ : full minus receiver-agnostic, with paired 95% CIs in the row below. Last row per dataset: A4 scored on YI , with its AUROC on YA in parentheses.
Figure 5: Behavioural history sharpens the receiver belief (960 audit histories per length). (a) Accuracy and (b) calibrated NLL by history length for genuine and shuffled history, with binomial 95% CIs in (a). (c) One-vs-rest AUROC at ∣Ht∣=5 against each receiver’s pooled YI , with 95% bootstrap CIs.
Calibrated, genuine history
Raw
Shuffled history
∣Ht∣
Accuracy ↑
NLL ↓
Brier ↓
ECE ↓
NLL ↓
Accuracy
NLL
0
20.0%
1.609
0.800
0.000
–
–
–
1
22.6%
1.634
0.811
0.056
1.626
20.9%
1.625
3
23.5%
1.604
0.798
0.013
1.603
18.4%
1.633
5
24.9%
1.593
0.794
0.014
1.594
21.3%
1.617
10
26.6%
1.591
0.793
0.031
1.593
19.9%
1.628
Appendix
Table 11: Receiver-type inference by history length ∣Ht∣ (960 audit histories per length). Length 0 is the uniform prior over five receivers, and “calibrated” uses temperature scaling fitted on the calibration split.
Dataset
Accuracy ↑
NLL ↓
Brier ↓
ECE ↓
FreebaseQA
21.5%
1.604
0.798
0.025
WebQSP
24.6%
1.598
0.796
0.028
GrailQA
20.5%
1.608
0.799
0.040
SQ-WD
18.3%
1.610
0.801
0.051
PopQA
25.9%
1.583
0.790
0.008
Mintaka
25.9%
1.588
0.792
0.017
Appendix
Table 12: Receiver-type inference at ∣Ht∣=5 : by dataset of the calibration tasks (top, 960 histories each) and row-normalised confusion matrix (bottom, %, true receiver in rows).
Figure 6: Risk-guided revision lowers measured interpretation failure on the 1,579 held-out questions. Paired differences in YI , ours minus each reference (points, 95% CIs, negative favours ours). (b) Per dataset against the original (filled) and the generic rewrite (hollow), post hoc.
Figure 7: One risk feature is shared by all receivers, and guidance removes it. (a) Learned scores srk and intercepts br (points of YI ). Boxed: the score and its 2.5% bootstrap quantile are both positive, the guidance rule for a known receiver. (b) YI per receiver. (c) Form of all sent messages (unchanged originals included) and YI per condition, with guided conditions in blue.
Condition
YI
YA
None
Distr.
Changed
Val. YI
Original message
4.16
59.33
0.93
3.20
0.0
5.01
Generic rewrite
3.85
58.15
0.88
2.95
97.2
4.97
Selection only
3.47
58.20
0.86
2.59
86.7
4.60
Posterior only
3.43
58.59
0.81
2.60
86.1
4.23
Ours
2.31
58.33
0.72
1.56
90.0
2.83
Guidance only
2.29
58.33
0.70
1.54
91.5
2.66
Appendix
Table 13: Experiment 4 results by condition. YI , YA , None , and Distr. (probes that choose the contrast task) are percentages on the 1,579 held-out questions, and Changed is the share of messages rewritten. Val. YI is on the 300 validation questions of the check before the main run.
YI (%)
Ours minus
Generic minus
Receiver
Original
Ours
original [95% CI]
generic
sel. only
post. only
original
Value
Nemotron
8.73
3.76
−4.96 [ −5.34 , −4.59 ]
−3.54
−2.70
−2.66
−1.42
−5.60
Gemma
2.84
2.16
−0.69 [ −1.04 , −0.35 ]
−0.88
−0.55
−0.42
+0.20
+0.07
gpt-oss
1.91
1.51
−0.40 [ −0.62 , −0.17 ]
−0.56
−0.63
−0.52
+0.16
+0.01
Mistral
5.27
2.65
−2.62 [ −3.00 , −2.25 ]
−2.04
−1.50
−1.60
−0.58
−4.09
Qwen
2.07
1.50
−0.57 [ −0.85 , −0.31 ]
−0.66
−0.40
−0.36
+0.09
−0.38
Appendix
Table 14: Experiment 4 results per receiver on the held-out questions. Ours minus each reference and the generic rewrite minus the original are paired differences in YI (points), and Value is ours minus the original on the 275 value-versus-attribute questions.
Figure 8: NetVoII allocates queries by decision value (primary cohort, interpretation objective, c=0.001 ). (a) YI against the share of episodes queried. (b) Paired differences at matched quotas with 95% CIs, equal to the net-utility contrasts with the sign reversed. (c) Share of episodes whose posterior mode is the true receiver. (d) Strict rule at six query costs c : measured YI (top) and share of episodes queried (bottom).
Figure 9: Information is not value (primary cohort). (a) Expected IG of the most informative query (nats) and share of episodes with positive predicted VoII, per dataset. (b) Share of episodes whose best query is worth more than a cost c , with pooled values marked at four of the tested costs.
Policy
10%
20%
30%
50%
Always
Primary cohort (test split). No query: 3.84 (0.9616), true identity: 3.77 (0.9623), adapted VoI † without queries: 4.42 (0.9558).
Random (3-seed mean)
3.84 (0.9615)
3.84 (0.9614)
3.83 (0.9614)
3.83 (0.9612)
3.84 (0.9606)
IG
3.83 (0.9616)
3.83 (0.9615)
3.83 (0.9614)
3.82 (0.9613)
3.80 (0.9610)
Adapted VoI †
4.42 (0.9557)
4.42 (0.9556)
4.42 (0.9555)
4.42 (0.9553)
–
NetVoII (ours)
3.80 (0.9619)
3.79 (0.9619)
3.78 (0.9619)
3.78 (0.9617)
–
Replication cohort. No query: 3.93 (0.9607), true identity: 3.91 (0.9609), adapted VoI † without queries: 4.27 (0.9573).
Appendix
Table 15: Experiment 5 budget grid (interpretation objective): measured YI (%) by exact query quota, with net utility at c=0.001 in parentheses. Bold: net-utility contrasts against both IG and random querying survive Holm correction. † Secondary comparator: its estimates never let a query change the message.
Cost c
Queried (%)
YI (%)
1−YA (%)
Net utility
Primary cohort (test split)
0
32.9
3.78
35.66
0.9622
0.0001
30.2
3.78
35.68
0.9622
0.0005
24.0
3.79
35.62
0.9620
0.001 (main)
18.2
3.79
35.68
0.9619
0.005
0.8
3.83
35.72
0.9617
Appendix
Table 16: Strict rule of Experiment 5 across query costs c (interpretation objective, net utility at each row’s cost). The last two rows of each block re-run IG and random querying at the query count of the c=0.001 rule.
YI at the 20% quota
Queried by NetVoII (%)
Unit
No query
NetVoII
Inf. gain
Random
20% quota
Strict
YI strict
True identity
FreebaseQA
2.16
2.16
2.15
2.16
1.5
1.1
2.16
2.10
WebQSP
5.07
5.07
5.07
5.07
0.0
0.0
5.07
5.14
GrailQA
6.96
6.96
6.96
6.96
0.0
0.0
6.96
6.96
SQ-WD
4.53
4.51
4.50
4.53
4.9
4.1
4.51
4.45
PopQA
2.61
2.42
2.62
2.60
49.8
45.1
2.42
2.52
Appendix
Table 17: Experiment 5 results per dataset and per receiver (primary cohort, interpretation objective): YI (%) at the pooled 20% quota and under the strict rule, with the share of episodes that NetVoII queries.
Figure 10: NetVoII spends a pooled quota where predicted value is high (primary cohort, interpretation objective). (a) Share of each dataset’s episodes queried at pooled quotas of 10–50% and under the strict rule. (b) Change in YI against no query at the 20% quota.
Multi-agent LLM systems rely on message passing among specialized agents to accomplish complex tasks. However, an upstream agent may provide useful information or an incorrect answer that causes a downstream agent to override a correct answer supported by its own evidence. Prior work has not clearly separated the benefits of communication from the damage caused by incorrect messages. We study this problem with controlled experiments across five benchmarks and five receivers, keeping the downstream task and evidence fixed while comparing answers under three conditions: no message, the upstream agent's original message, or a message with the opposite conclusion. Our experiments reveal three key findings. First, messages often help when the downstream agent would otherwise answer incorrectly. Second, messages can also hurt: when the downstream agent would answer correctly without a message, an incorrect upstream message changes the answer in up to 32% of cases. Third, in 94% of audited harmful cases, the downstream agent copies the upstream's specific wrong answer--a pattern we term answer substitution. Removing unreliable messages recovers part of the lost accuracy, suggesting that communication should be selective based on upstream reliability and the evidence already available to the downstream agent.
Yaxin Gong, Gangyi Zhang, Chongming Gao +7
University of Science and Technology of China · Qwen Business Unit of Alibaba · National University of Singapore
Large Language Models (LLMs) have enabled collaborative Multi-Agent (MA) systems, where interacting agents improve performance through diverse reasoning and iterative refinement. However, these systems remain vulnerable to error propagation, where early-stage information degrades downstream reasoning. To address this, we conduct a systematic analysis of inter-agent communication to identify which information drives MA performance. We find that the absence of reasoning and verification in inter-agent communication significantly degrades performance. Based on these insights, we propose Category-Aware Recovery Augmentation (technique), which enforces the presence of critical information during communication. recovers up to 86.2% of failed cases. Our results highlight the key role of information quality in effective MA collaboration. Our code is available at https://anonymous.4open.science/r/cara_mas
Yong Jin Chun, Iftekhar Ahmed
Department of Informatics University of California, Irvine Irvine, CA, 92618
LLMs are increasingly deployed as post-hoc explainers of AI-generated outputs, yet it remains unclear whether they can reliably communicate probabilistic information in natural language. For this role to be viable, models must produce identical verbal descriptions for identical inputs, and select descriptions that accurately reflect the magnitude of the underlying numerical quantities. We evaluate whether nine LLMs meet these requirements within a two-stage prediction pipeline, in which an upstream model has produced probabilistic outputs characterized by their likelihood and uncertainty, and LLMs are tasked with selecting an appropriate verbal descriptor for each. We simulate predictions from an upstream model by taking samples from a Beta distribution parameterized by its mode and prior sample size. We then prompt LLMs to explain these predictions under six domain contexts and with ten temperature settings, and repeating each experiment ten times. We find that LLMs are generally consistent but miscalibrated, with substantially weaker performance on uncertainty than on likelihood tasks. Providing models with precomputed summary statistics (mode and prior sample size) reduced sensitivity to contextual framing but did not resolve the underlying miscalibration, suggesting that the bottleneck resides in the verbalization step itself. These findings indicate that current LLMs do not yet constitute reliable zero-shot standalone risk communication tools for probabilistic predictions.
Diego Cerda-Mardini, Sarath Chandar, Sreenath Madathil
Faculty of Dental Medicine and Oral Health Sciences, McGill University, Montréal, QC, Canada · Chandar Research Lab, Polytechnique Montréal, Montréal, H3T 1J4, Canada · Mila – Québec Artificial Intelligence Institute, Montréal, H2S 3H1, Canada +2